International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1704964

1704964 Vol 7 · Issue 2 Download Paper

Classification Preservation Using Assorted Dimensionality Reduction Techniques

Usman A. Baba Augustine S. Nsang

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

Abstract

In this paper, we implement the perceptron classification algorithm and apply it to three two-class datasets which include the student, weather and ionosphere datasets. Then the k-Nearest Neighbors classification algorithm is also applied to the same two-class datasets. Each dataset is then reduced using fourteen different dimensionality reduction techniques. The perceptron and k-nearest neighbor classification algorithms are then applied to each reduced set and the performances of the dimensionality reduction techniques in preserving the classification of a dataset by the k-nearest neighbors and perceptron classification algorithm are compared. The extent to which the classification of a dataset is preserved by a given dimensionality reduction technique is evaluated using the rand index and confusion matrices.

Keywords

Classification, Confusion Matrix, Dimensionality Reduction, Eager Learner, k-Nearest Neighbors, Lazy Learner, Perceptron, Rand Index

References

[1] N. Sharma and K. Saroha, “Study of dimension reduction methodologies in data mining,” in International Conference on Computing, Communication and Automation, 2015, pp. 133–137.

[2] I. K. Fodor, “A survey of dimension reduction techniques,” Center for Applied Scientific Computing, Lawrence Livermore National Laboratory, no. 1, pp. 1–18, 2002.

[3] D. H. Deshmukh, T. Ghorpade, and P. Padiya, “Improving classification using preprocessing and machine learning algorithms on NSL-KDD dataset,” in Proceedings - 2015 International Conference on Communication, Information and Computing Technology, ICCICT 2015, 2015.

[4] A. S. Nsang, I. Diaz, and A. Ralescu, “Ensemble Clustering based on Heterogeneous Dimensionality Reduction Methods and Context-dependent Similarity Measures,” Int. J. Adv. Sci. Technol., vol. 64, pp. 101–118, 2014.

[5] A. S. Nsang , F. Oguntoyinbo, H. Yusuf, and A. Maikori, “A New Random Approach To Dimensionality Reduction, in "Int’l Conf. on Advances in Big Data Analytics", pp. 69 – 74, 2015.

[6] I. Kavakiotis, O. Tsave, A. Salifoglou, N. Maglaveras, I. Vlahavas, and I. Chouvarda, “Machine Learning and Data Mining Methods in Diabetes Research,” Comput. Struct. Biotechnol. J., vol. 15, pp. 104–116, 2017.

[7] T. M. Mitchell, Machine Learning, vol. 1, no. 3. 1997.

[8] S. B. Kotsiantis, “Supervised machine learning: A review of classification techniques,” Informatica, vol. 31, pp. 249–268, 2007.

[9] S. B. Kotsiantis, I. D. Zaharakis, and P. E. Pintelas, “Machine learning: A review of classification and combining techniques,” Artif. Intell. Rev., vol. 26, no. 3, pp. 159–190, 2006.

[10] M. Capó, A. Pérez, and J. A. Lozano, “An efficient approximation to the K-means clustering for massive data,” Knowledge-Based Systems, 2016.

[11] Y. H. and W. Lam, “Lazy Learning for Classication Based on Query Projections,” in Proceedings of the 2005 SIAM International Conference on Data Mining, 2005, pp. 227–238.

[12] N. Singh, “Malware Analysis , Clustering and Classification : A Literature Review,” IJCST Int. J. Comput. Sci. Technol., vol. 8491, pp. 68–72, 2015.

[13] I. M. Galván, J. M. Valls, M. García, and P. Isasi, “A lazy learning approach for building classification models,” Int. J. Intell. Syst., vol. 26, no. 8, pp. 773–786, 2011.

[14] A. S. Nsang, A. M. Bello, and H. Shamsudeen, “Image Reduction Using Assorted Dimensionality Reduction Techniques,” in Proceedings of the 26th Modern Artificial Intelligence and Cognitive Science Conference, pp 139 - 146, 2015.

[15] E. Bingham and H. Mannila, “Random projection in dimensionality reduction: Applications To Image And Text Data,” Int. Conf. Knowl. Discov. Data Min., pp. 245–250, 2001.

[16] Augustine S. Nsang and Anca Ralescu. Approaches to Dimensionality Reduction to a Subset of the Original Dimensions. In Proceedings of the Twenty-First Midwest Artificial Intelligence and Cognitive Science Conference, 70-77, 2010.

[17] Augustine Nsang. Novel Approaches to Dimensionality Reduction and Applications: An Empirical Study. Lambert Academic Publishing, Saarbrücken, Germany, 2011.

[18] A. S. Nsang, D. Edi, and C. Ahanonu, “Query-Based Dimensionality Reduction Applied To Images,” in Int’l Conf. on Advances in Big Data Analytics, 2015, no. 2, pp. 81–86.

[19] L. E. Peterson, “K-nearest neighbors,” Scholarpedia, vol. 4, no. 2, p. 1883, 2009.

[20] W. Ertel, Introduction to Artificial Intelligence. 2011.

[21] L.-Y. Hu, M.-W. Huang, S.-W. Ke, and C.-F. Tsai, “The distance function effect on k-nearest neighbor classification for medical datasets,” Springerplus, vol. 5, no. 1, p. 1304, 2016.

[22] F. Rosenblatt, “The perceptron: a probabilistic model for information storage and organization in the brain.,” Psychol. Rev., vol. 65, no. 6, pp. 386–408, 1958.

[23] S. Haykin, Neural Networks and Learning Machines, vol. 3. 2008.

[24] S. Haykin, “Rosenblatt’s Perceptron,” Neural Networks Learn. Mach., no. 1943, pp. 47–67, 2009.

[25] K. Bache and M. Lichman, “UCI Machine Learning Repository,” University of California Irvine School of Information, vol. 2008, no. 14/8. p. 0, 2013.

[26] J. M. Santos and M. Embrechts, “On the Use of the Adjusted Rand Index as a Metric for Evaluating Supervised Classification.pdf,” in 19th International Conference on Artificial Neural Networks, 2009, pp. 1–10.

[27] S. Visa, B. Ramsay, A. Ralescu, and E. Van Der Knaap, “Confusion matrix-based feature selection,” in CEUR Workshop Proceedings, 2011, vol. 710, pp. 120–127.

[28] S. Singh and R. Singla, “Comparative Performance of Fault-Prone Prediction Classes with K-means Clustering and MLP,” in Proceedings of the Second International Conference on Information and Communication Technology for Competitive Strategies, 2016.

[29] M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, pp. 427–437, 2009.

[30] U. A. Baba, A. S. Nsang and O. Adeseye. “Three Novel Approaches to Dimensionality Reduction.” In Proceedings of the International Conference of Artificial Intelligence, Las Vegas, 2018.

How to cite this paper

Usman A. Baba, Augustine S. Nsang "Classification Preservation Using Assorted Dimensionality Reduction Techniques" Iconic Research And Engineering Journals Volume 7 Issue 2 2023 Page 245-257
Usman A. Baba, Augustine S. Nsang "Classification Preservation Using Assorted Dimensionality Reduction Techniques" Iconic Research And Engineering Journals, vol. 7, no. 2, Aug. 2023
Usman A. Baba, Augustine S. Nsang (2023). Classification Preservation Using Assorted Dimensionality Reduction Techniques. Iconic Research And Engineering Journals, 7(2).
Usman A. Baba, Augustine S. Nsang "Classification Preservation Using Assorted Dimensionality Reduction Techniques" Iconic Research And Engineering Journals, vol. 7, no. 2, Aug. 2023.
@article{1704964,
      author = {Usman A. Baba, Augustine S. Nsang},
      title = {Classification Preservation Using Assorted Dimensionality Reduction Techniques},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {7},
      number = {2},
      pages = {245-257},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1704964.pdf},
      abstract = {In this paper, we implement the perceptron classification algorithm and apply it to three two-class datasets which include the student, weather and ionosphere datasets. Then the k-Nearest Neighbors classification algorithm is also applied to the same two-class datasets. Each dataset is then reduced using fourteen different dimensionality reduction techniques. The perceptron and k-nearest neighbor classification algorithms are then applied to each reduced set and the performances of the dimensionality reduction techniques in preserving the classification of a dataset by the k-nearest neighbors and perceptron classification algorithm are compared. The extent to which the classification of a dataset is preserved by a given dimensionality reduction technique is evaluated using the rand index and confusion matrices.},
      keywords = {Classification, Confusion Matrix, Dimensionality Reduction, Eager Learner, k-Nearest Neighbors, Lazy Learner, Perceptron, Rand Index},
      month = {August},
  }