Home / Current Issue / Paper 1718411
Machine Learning for Clinical Decision Support: A Review of Feature Selection in Disease Prediction
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
DOI: https://doi.org/10.64388/IREV5I1-1718411
Abstract
Machine learning-based clinical decision support systems offer transformative potential for improving diagnostic accuracy and risk stratification, but adoption remains constrained by the challenge of identifying minimal clinically meaningful feature sets achieving adequate predictive performance while supporting interpretability and reducing data collection burden. This paper presents a comprehensive review of feature selection approaches applied to disease prediction machine learning models, examining filter methods, wrapper methods, embedded methods, and evolutionary optimisation including genetic algorithms, with systematic evaluation across liver disease, cardiovascular disease, diabetes, and oncological risk stratification domains. A consolidated best-practice framework integrating complementary filter and wrapper approaches within a cross-validated evaluation protocol is proposed. Genetic algorithm feature selection on the UCI Indian Liver Patient Dataset achieved six-feature XGBoost AUC-ROC of 0.841 with 40 percent feature count reduction. Comparative tables of feature selection methods and NLP approaches are included.
Keywords
Feature Selection, Clinical Decision Support, Machine Learning, Genetic Algorithm, Disease Prediction, SHAP, Liver Disease, Ensemble Methods, Missing Data, Overfitting
How to cite this paper
@article{1718411,
author = {Maryann Inimfon Atakpa, Toyosi Abolaji},
title = {Machine Learning for Clinical Decision Support: A Review of Feature Selection in Disease Prediction},
journal = {Iconic Research And Engineering Journals},
year = {2021},
volume = {5},
number = {1},
pages = {570-593},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718411.pdf},
abstract = {Machine learning-based clinical decision support systems offer transformative potential for improving diagnostic accuracy and risk stratification, but adoption remains constrained by the challenge of identifying minimal clinically meaningful feature sets achieving adequate predictive performance while supporting interpretability and reducing data collection burden. This paper presents a comprehensive review of feature selection approaches applied to disease prediction machine learning models, examining filter methods, wrapper methods, embedded methods, and evolutionary optimisation including genetic algorithms, with systematic evaluation across liver disease, cardiovascular disease, diabetes, and oncological risk stratification domains. A consolidated best-practice framework integrating complementary filter and wrapper approaches within a cross-validated evaluation protocol is proposed. Genetic algorithm feature selection on the UCI Indian Liver Patient Dataset achieved six-feature XGBoost AUC-ROC of 0.841 with 40 percent feature count reduction. Comparative tables of feature selection methods and NLP approaches are included.},
keywords = {Feature Selection, Clinical Decision Support, Machine Learning, Genetic Algorithm, Disease Prediction, SHAP, Liver Disease, Ensemble Methods, Missing Data, Overfitting},
month = {July},
doi = {https://doi.org/10.64388/IREV5I1-1718411}
}