Home / Current Issue / Paper 1700077
A New Approach For Classification Algorithm In Data Mining
Subject area: Science, Engineering and Technology · Area of research: Data Mining
Abstract
Data mining is the process of analyzing hidden patterns of data according to different perspectives for categorization into useful information, which is collected and assembled in common areas, such as data warehouses, for efficient analysis, data mining algorithms, facilitating business decision making and other information requirements to ultimately cut costs and increase revenue. I consider classification techniques that are based on statistical and AI techniques to perform controlled experiment data characteristics are systematically altered to introduce imperfections such as nonlinearity, unequal covariance. Two machine learning algorithms Naive Bayes and Support Vector Machine (SVM) is used to build models for the automatic classification of the tweets, and these models were evaluated across the metrics of accuracy, precision, recall, area under curve and measure. The results reveal that the proposed sampling strategy makes more judicious use of data points by selecting locations that clarify high level structures in data, rather than choosing points that merely improve quality of function approximation.
Keywords
Data Mining, Knowledge discovery data base, Decision tree, Extreme learning machine, Support Vector Machine
References
[1] M. Gaviano, D.E. Kvasov, D. Lera, and Y.D. Sergeyev. Algorithm 829: Software for Generation of Classes of Test Functions with Known Local and Global Minima for Global Optimization. ACM Transactions on Mathematical Software, Vol. 29(4): pages 469– 480, Dec 2003.
[2] N. Cristianini and J. Shawe-Taylor. An Introduction to Support Vector Machines and Other Kernel-Based Learning Methods. Cambridge University Press, 2000.
[3] C. Bailey-Kellogg, F. Zhao, and K. Yip. Spatial Aggregation: Language and Applications. In Proc. AAAI, pages 517–522, 1996.
[4] K. Brinker. Incorporating Diversity in Active Learning with Support Vector Machines. In Proceedings of the Twentieth International Conference on Machine Learning (ICML’03), pages 59–66, 2003.
[5] D.A. Cohn, Z. Ghahramani, and M.I. Jordan. Active Learning with Statistical Models. Journal of Artificial Intelligence Research, Vol. 4: pages 129–145, 1996.
[6] D. Cornford, I.T. Nabney, and C.K.I. Williams. Adding Constrained Discontinuities to Gaussian Process Models of Wind Fields. In Proceedings of NIPS, pages 861–867, 1998.
[7] C. Bailey-Kellogg and F. Zhao. Influence-Based Model Decomposition for Reasoning about Spatially Distributed Physical Systems. Artificial Intelligence, Vol. 130(2): pages 125–166, 2001
[8] I.T. Nabney. Netlab: Algorithms for Pattern Recognition. Springer-Verlag, 2002.
[9] R.G. Easterling. Comment on ‘Design and Analysis of Computer Experiments’. Statistical Science, 4(4):425– 427, 1989.
[10] J. Garcke and M. Griebel. Data Mining with Sparse Grids using Simplicial Basis Functions. In Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 87–96, 2001.
[11] C.Bailey-Kellogg and N. Ramakrishnan. AmbiguityDirected Sampling for Qualitative Analysis of Sparse Data from Spatially Distributed Physical Systems. In Proc. IJCAI, pages 43–50, 2001.
[12] M. Gaviano, D.E. Kvasov, D. Lera, and Y.D. Sergeyev. Algorithm 829: Software for Generation of Classes of Test Functions with Known Local and Global Minima for Global Optimization. ACM Transactions on Mathematical Software, Vol. 29(4): pages 469– 480, Dec 2003.
[13] X. Huang and F. Zhao. Relation-Based Aggregation: Finding Objects in Large Spatial Datasets. In Proceedings of the 3rd International Symposium on Intelligent Data Analysis, 1999.
[14] J.-N. Hwang, J.J. Choi, S. Oh, and R.J. Marks II. Query-based Learning Applied to Partially Trained Multilayer Perceptrons. IEEE Transactions on Neural Networks, Vol. 2(1): pages 131–136, 1991.
[15] A.G. Journel and C.J. Huijbregts. Mining Geostatistics. Academic Press, New York, 1992.
[16] S. Tong and D. Koller. Support Vector Machine Active Learning with Applications to Text Classification. Journal of Machine Learning Research, Vol. 2: pages 45–66, 2001.
[17] J. Koehler and A. Owen. Computer Experiments. In S. Ghosh and C. Rao, editors, Handbook of Statistics: Design and Analysis of Experiments, pages 261–308. North Holland, 1996.
[18] D.J. MacKay. Information-Based Objective Functions for Active Data Selection. Neural Computation, Vol. 4(4): pages 590–604, 1992.
[19] C. Currin, T. Mitchell, M. Morris, and D. Ylvisaker. Bayesian Prediction of Deterministic Functions, with Applications to the Design and Analysis of Computer Experiments. J. Amer. Stat. Assoc., Vol. 86: pages 953– 963, 1991.
[20] R.M. Neal. Monte Carlo Implementations of Gaussian Process Models for Bayesian Regression and Classification. Technical Report 9702, Department of Statistics, University of Toronto, Jan 1997.
[21] R.T. Ng and J. Han. CLARANS: A Method for Clustering Objects for Spatial Data Mining. IEEE Transactions on Knowledge and Data Engineering, Vol. 14(5): pages 1003–1016, 2002.
[22] I. Ord´onez ˜ and F. Zhao. STA: Spatio-Temporal Aggregation with Applications to Analysis of DiffusionReaction Phenomena. In Proc. AAAI, pages 517–523, 2000.
[23] N. Ramakrishnan and C. Bailey-Kellogg. Sampling Strategies for Mining in Data-Scarce Domains. IEEE/AIP CiSE, Vol. 4(4): pages 31– 43, 2002.
[24] N. Ramakrishnan and C. Bailey-Kellogg. Gaussian Process Models of Spatial Aggregation Algorithms. In Proc. IJCAI, pages 1045–1051, 2003.
[25] J. Sacks, W.J. Welch, T.J. Mitchell, and H.P. Wynn. Design and Analysis of Computer Experiments. Statistical Science, Vol. 4(4): pages 409–435, 1989.
[26] G. Karypis, E.-H. Han, and V. Kumar. Chameleon: Hierarchical Clustering using Dynamic Modeling. IEEE Computer, Vol. 32(8): pages 68–75, 1999.
How to cite this paper
@article{1700077,
author = {M. Thillaikarasi},
title = {A New Approach For Classification Algorithm In Data Mining},
journal = {Iconic Research And Engineering Journals},
year = {2017},
volume = {1},
number = {4},
pages = {77-82},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1700077.pdf},
abstract = {Data mining is the process of analyzing hidden patterns of data according to different perspectives for categorization into useful information, which is collected and assembled in common areas, such as data warehouses, for efficient analysis, data mining algorithms, facilitating business decision making and other information requirements to ultimately cut costs and increase revenue. I consider classification techniques that are based on statistical and AI techniques to perform controlled experiment data characteristics are systematically altered to introduce imperfections such as nonlinearity, unequal covariance. Two machine learning algorithms Naive Bayes and Support Vector Machine (SVM) is used to build models for the automatic classification of the tweets, and these models were evaluated across the metrics of accuracy, precision, recall, area under curve and measure. The results reveal that the proposed sampling strategy makes more judicious use of data points by selecting locations that clarify high level structures in data, rather than choosing points that merely improve quality of function approximation.},
keywords = {Data Mining, Knowledge discovery data base, Decision tree, Extreme learning machine, Support Vector Machine},
month = {October},
}