International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722828

1722828 Vol 10 · Issue 3 Download Paper

Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data

Patrick Nwogu Prof. Chigozirim Ajaegu Dr. Faruk Umar Ambursa Dr. Femi Adeluyi

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

DOI: https://doi.org/10.64388/IREV10I3-1722828

Abstract

Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.

Keywords

dyslexia; learning disability; machine learning; Random Forest; XGBoost; Extra Trees; RF-RFE; ensemble learning; educational data mining.

How to cite this paper

Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data" Iconic Research And Engineering Journals Volume 10 Issue 3 2026 Page 531-533 https://doi.org/10.64388/IREV10I3-1722828
Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026, doi: https://doi.org/10.64388/IREV10I3-1722828
Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi (2026). Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data. Iconic Research And Engineering Journals, 10(3). doi: https://doi.org/10.64388/IREV10I3-1722828
Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026. Crossref, https://doi.org/10.64388/IREV10I3-1722828
@article{1722828,
      author = {Patrick Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi},
      title = {Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {3},
      pages = {531-533},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722828.pdf},
      abstract = {Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.},
      keywords = {dyslexia; learning disability; machine learning; Random Forest; XGBoost; Extra Trees; RF-RFE; ensemble learning; educational data mining.},
      month = {September},
      doi = {https://doi.org/10.64388/IREV10I3-1722828}
  }