Home / Current Issue / Paper 1715172
Empirical Evaluation of Learning Curve Cross-Validation for Efficient Model Selection
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
Abstract
Model selection in machine learning is commonly performed using cross-validation, where candidate models are evaluated to estimate their generalization performance. Although reliable, this approach becomes computationally expensive when many models must be evaluated repeatedly on large datasets. Learning Curve Cross-Validation (LCCV) addresses this issue by evaluating models on progressively larger subsets of data and pruning weak candidates early. In this work, we implement the LCCV algorithm and evaluate it on several classification tasks. The implementation estimates performance at different training sizes using repeated cross-validation with confidence intervals. Models are pruned using optimistic learning curve extrapolation, while the Morgan–Mercer–Flodin (MMF) model is used to skip intermediate evaluation points. Experiments on real-world and synthetic datasets compare LCCV with traditional full cross-validation in terms of runtime, model selection agreement, and pruning behavior. Results show that LCCV can prune many candidate models and significantly reduce runtime on larger datasets, while introducing some overhead on smaller datasets.
Keywords
Learning Curve Cross-Validation, Model Selection, Learning Curves, Early Pruning, Cross-Validation
References
[1] A. Klein, S. Falkner, J. T. Springenberg, and F. Hutter, “Fast and Informative Model Selection using Learning Curve Cross-Validation,” Proceedings of the International Conference on Learning Representations (ICLR), 2017.
[2] K. Jamieson and A. Talwalkar, “Non-stochastic Best Arm Identification and Hyperparameter Optimization,” Proceedings of the International Conference on Artificial Intelligence and Statistics (AISTATS), 2016.
[3] T. Domhan, J. T. Springenberg, and F. Hutter, “Speeding Up Automatic Hyperparameter Optimization of Deep Neural Networks by Extrapolation of Learning Curves,” Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2015.
[4] J. Bergstra and Y. Bengio, “Random Search for Hyper-Parameter Optimization,” Journal of Machine Learning Research, vol. 13, pp. 281–305, 2012.
[5] J. N. van Rijn, S. Bischl, and J. Vanschoren, “Fast Algorithm Selection Using Learning Curves,” Proceedings of the International Symposium on Intelligent Data Analysis, 2015.
[6] J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian Optimization of Machine Learning Algorithms,” Advances in Neural Information Processing Systems (NeurIPS), 2012.
[7] F. Hutter, H. H. Hoos, and K. Leyton-Brown,“Sequential Model-Based Optimization for General Algorithm Configuration,” Proceedings of the International Conference on Learning and Intelligent Optimization (LION), 2011.
[8] L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar, “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization,” Journal of Machine Learning Research, vol. 18, no. 185, pp. 1–52, 2017.
[9] C. Thornton, F. Hutter, H. H. Hoos, and K. Leyton-Brown, “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms,” Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2013.
[10] M. Feurer, A. Klein, K. Eggensperger, J. T. Springenberg, M. Blum, and F. Hutter,“Auto-sklearn: Efficient and Robust Automated Machine Learning,” Advances in Neural Information Processing Systems (NeurIPS), 2015.
[11] A. Klein, S. Falkner, N. Mansur, and F. Hutter,“Learning Curve Prediction with Bayesian Neural Networks,” International Conference on Learning Representations (ICLR), 2017.
[12] R. Kohavi, “A Study of Cross-Validation and Bootstrap for Accuracy Estimation and Model Selection,” Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 1995.
[13] C. M. Bishop, Pattern Recognition and Machine Learning. New York: Springer, 2006.
How to cite this paper
@article{1715172,
author = {Dr. M. Pompapathi, N. Siva Parvathi, V. Pavani, S. Varshini, V. Manikanta},
title = {Empirical Evaluation of Learning Curve Cross-Validation for Efficient Model Selection},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {9},
pages = {1243-1251},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1715172.pdf},
abstract = {Model selection in machine learning is commonly performed using cross-validation, where candidate models are evaluated to estimate their generalization performance. Although reliable, this approach becomes computationally expensive when many models must be evaluated repeatedly on large datasets. Learning Curve Cross-Validation (LCCV) addresses this issue by evaluating models on progressively larger subsets of data and pruning weak candidates early. In this work, we implement the LCCV algorithm and evaluate it on several classification tasks. The implementation estimates performance at different training sizes using repeated cross-validation with confidence intervals. Models are pruned using optimistic learning curve extrapolation, while the Morgan–Mercer–Flodin (MMF) model is used to skip intermediate evaluation points. Experiments on real-world and synthetic datasets compare LCCV with traditional full cross-validation in terms of runtime, model selection agreement, and pruning behavior. Results show that LCCV can prune many candidate models and significantly reduce runtime on larger datasets, while introducing some overhead on smaller datasets.},
keywords = {Learning Curve Cross-Validation, Model Selection, Learning Curves, Early Pruning, Cross-Validation},
month = {March},
doi = {https://doi.org/10.64388/IREV9I9-1715172}
}