Home / Current Issue / Paper 1704964
Classification Preservation Using Assorted Dimensionality Reduction Techniques
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
In this paper, we implement the perceptron classification algorithm and apply it to three two-class datasets which include the student, weather and ionosphere datasets. Then the k-Nearest Neighbors classification algorithm is also applied to the same two-class datasets. Each dataset is then reduced using fourteen different dimensionality reduction techniques. The perceptron and k-nearest neighbor classification algorithms are then applied to each reduced set and the performances of the dimensionality reduction techniques in preserving the classification of a dataset by the k-nearest neighbors and perceptron classification algorithm are compared. The extent to which the classification of a dataset is preserved by a given dimensionality reduction technique is evaluated using the rand index and confusion matrices.
Keywords
Classification, Confusion Matrix, Dimensionality Reduction, Eager Learner, k-Nearest Neighbors, Lazy Learner, Perceptron, Rand Index
References
[1] N. Sharma and K. Saroha, “Study of dimension reduction methodologies in data mining,” in International Conference on Computing, Communication and Automation, 2015, pp. 133–137.
[2] I. K. Fodor, “A survey of dimension reduction techniques,” Center for Applied Scientific Computing, Lawrence Livermore National Laboratory, no. 1, pp. 1–18, 2002.
[3] D. H. Deshmukh, T. Ghorpade, and P. Padiya, “Improving classification using preprocessing and machine learning algorithms on NSL-KDD dataset,” in Proceedings - 2015 International Conference on Communication, Information and Computing Technology, ICCICT 2015, 2015.
[4] A. S. Nsang, I. Diaz, and A. Ralescu, “Ensemble Clustering based on Heterogeneous Dimensionality Reduction Methods and Context-dependent Similarity Measures,” Int. J. Adv. Sci. Technol., vol. 64, pp. 101–118, 2014.
[5] A. S. Nsang , F. Oguntoyinbo, H. Yusuf, and A. Maikori, “A New Random Approach To Dimensionality Reduction, in "Int’l Conf. on Advances in Big Data Analytics", pp. 69 – 74, 2015.
[6] I. Kavakiotis, O. Tsave, A. Salifoglou, N. Maglaveras, I. Vlahavas, and I. Chouvarda, “Machine Learning and Data Mining Methods in Diabetes Research,” Comput. Struct. Biotechnol. J., vol. 15, pp. 104–116, 2017.
[7] T. M. Mitchell, Machine Learning, vol. 1, no. 3. 1997.
[8] S. B. Kotsiantis, “Supervised machine learning: A review of classification techniques,” Informatica, vol. 31, pp. 249–268, 2007.
[9] S. B. Kotsiantis, I. D. Zaharakis, and P. E. Pintelas, “Machine learning: A review of classification and combining techniques,” Artif. Intell. Rev., vol. 26, no. 3, pp. 159–190, 2006.
[10] M. Capó, A. Pérez, and J. A. Lozano, “An efficient approximation to the K-means clustering for massive data,” Knowledge-Based Systems, 2016.
[11] Y. H. and W. Lam, “Lazy Learning for Classication Based on Query Projections,” in Proceedings of the 2005 SIAM International Conference on Data Mining, 2005, pp. 227–238.
[12] N. Singh, “Malware Analysis , Clustering and Classification : A Literature Review,” IJCST Int. J. Comput. Sci. Technol., vol. 8491, pp. 68–72, 2015.
[13] I. M. Galván, J. M. Valls, M. García, and P. Isasi, “A lazy learning approach for building classification models,” Int. J. Intell. Syst., vol. 26, no. 8, pp. 773–786, 2011.
[14] A. S. Nsang, A. M. Bello, and H. Shamsudeen, “Image Reduction Using Assorted Dimensionality Reduction Techniques,” in Proceedings of the 26th Modern Artificial Intelligence and Cognitive Science Conference, pp 139 - 146, 2015.
[15] E. Bingham and H. Mannila, “Random projection in dimensionality reduction: Applications To Image And Text Data,” Int. Conf. Knowl. Discov. Data Min., pp. 245–250, 2001.
[16] Augustine S. Nsang and Anca Ralescu. Approaches to Dimensionality Reduction to a Subset of the Original Dimensions. In Proceedings of the Twenty-First Midwest Artificial Intelligence and Cognitive Science Conference, 70-77, 2010.
[17] Augustine Nsang. Novel Approaches to Dimensionality Reduction and Applications: An Empirical Study. Lambert Academic Publishing, Saarbrücken, Germany, 2011.
[18] A. S. Nsang, D. Edi, and C. Ahanonu, “Query-Based Dimensionality Reduction Applied To Images,” in Int’l Conf. on Advances in Big Data Analytics, 2015, no. 2, pp. 81–86.
[19] L. E. Peterson, “K-nearest neighbors,” Scholarpedia, vol. 4, no. 2, p. 1883, 2009.
[20] W. Ertel, Introduction to Artificial Intelligence. 2011.
[21] L.-Y. Hu, M.-W. Huang, S.-W. Ke, and C.-F. Tsai, “The distance function effect on k-nearest neighbor classification for medical datasets,” Springerplus, vol. 5, no. 1, p. 1304, 2016.
[22] F. Rosenblatt, “The perceptron: a probabilistic model for information storage and organization in the brain.,” Psychol. Rev., vol. 65, no. 6, pp. 386–408, 1958.
[23] S. Haykin, Neural Networks and Learning Machines, vol. 3. 2008.
[24] S. Haykin, “Rosenblatt’s Perceptron,” Neural Networks Learn. Mach., no. 1943, pp. 47–67, 2009.
[25] K. Bache and M. Lichman, “UCI Machine Learning Repository,” University of California Irvine School of Information, vol. 2008, no. 14/8. p. 0, 2013.
[26] J. M. Santos and M. Embrechts, “On the Use of the Adjusted Rand Index as a Metric for Evaluating Supervised Classification.pdf,” in 19th International Conference on Artificial Neural Networks, 2009, pp. 1–10.
[27] S. Visa, B. Ramsay, A. Ralescu, and E. Van Der Knaap, “Confusion matrix-based feature selection,” in CEUR Workshop Proceedings, 2011, vol. 710, pp. 120–127.
[28] S. Singh and R. Singla, “Comparative Performance of Fault-Prone Prediction Classes with K-means Clustering and MLP,” in Proceedings of the Second International Conference on Information and Communication Technology for Competitive Strategies, 2016.
[29] M. Sokolova and G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, pp. 427–437, 2009.
[30] U. A. Baba, A. S. Nsang and O. Adeseye. “Three Novel Approaches to Dimensionality Reduction.” In Proceedings of the International Conference of Artificial Intelligence, Las Vegas, 2018.
How to cite this paper
@article{1704964,
author = {Usman A. Baba, Augustine S. Nsang},
title = {Classification Preservation Using Assorted Dimensionality Reduction Techniques},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {2},
pages = {245-257},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1704964.pdf},
abstract = {In this paper, we implement the perceptron classification algorithm and apply it to three two-class datasets which include the student, weather and ionosphere datasets. Then the k-Nearest Neighbors classification algorithm is also applied to the same two-class datasets. Each dataset is then reduced using fourteen different dimensionality reduction techniques. The perceptron and k-nearest neighbor classification algorithms are then applied to each reduced set and the performances of the dimensionality reduction techniques in preserving the classification of a dataset by the k-nearest neighbors and perceptron classification algorithm are compared. The extent to which the classification of a dataset is preserved by a given dimensionality reduction technique is evaluated using the rand index and confusion matrices.},
keywords = {Classification, Confusion Matrix, Dimensionality Reduction, Eager Learner, k-Nearest Neighbors, Lazy Learner, Perceptron, Rand Index},
month = {August},
}