Home / Current Issue / Paper 1722828
Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
DOI: https://doi.org/10.64388/IREV10I3-1722828
Abstract
Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.
Keywords
dyslexia; learning disability; machine learning; Random Forest; XGBoost; Extra Trees; RF-RFE; ensemble learning; educational data mining.
How to cite this paper
@article{1722828,
author = {Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi},
title = {Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {3},
pages = {531-533},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722828.pdf},
abstract = {Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.},
keywords = {dyslexia; learning disability; machine learning; Random Forest; XGBoost; Extra Trees; RF-RFE; ensemble learning; educational data mining.},
month = {September},
doi = {https://doi.org/10.64388/IREV10I3-1722828}
}