International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722829

1722829 Vol 10 · Issue 3 Download Paper

Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning

Patrick Ufomba Nwogu Prof. Chigozirim Ajaegu Dr. Faruk Umar Ambursa Dr. Femi Adeluyi

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning

DOI: https://doi.org/10.64388/IREV10I3-1722829

Abstract

This study develops and empirically evaluates an ensemble machine-learning framework for dyslexia detection using Random-Forest Recursive Feature Elimination (RF-RFE), Random Forest (RF), XGBoost, Extra Trees (ET), and probability-level logistic-regression stacking. The analysis uses the supplied Dyt-desktop.csv dataset, which contains 3,644 observations, 196 predictors and one binary dyslexia outcome. The target distribution is imbalanced: 392 dyslexia cases (10.76%) and 3,252 non-dyslexia cases (89.24%). Five-fold stratified out-of-fold evaluation was used. Class-weighted RF, XGBoost and ET were compared with RF-RFE variants, while the proposed ensemble combined out-of-fold probability predictions from RF, XGBoost and ET through logistic regression. At the default 0.50 threshold, XGBoost produced the strongest individual accuracy (90.64%), while RF-RFE + XGBoost achieved 90.81% accuracy, 64.77% precision, 31.89% recall, 97.91% specificity, F1=0.4274, ROC-AUC=0.8858 and MCC=0.4122. The proposed stacking ensemble produced ROC-AUC=0.8839, recall=71.17%, F1=0.5171 and MCC=0.4644, but lower accuracy (85.70%) because its class-weighted meta-classifier generated substantially more positive predictions. These findings demonstrate that model ranking changes according to the evaluation objective: RF-RFE + XGBoost is strongest on accuracy and ROC-AUC, whereas the proposed ensemble provides the strongest minority-class recall, F1 and MCC at the tested threshold. The results support threshold optimization and precision-recall analysis as essential components of dyslexia screening.

Keywords

dyslexia; machine learning; ensemble learning; RF-RFE; Random Forest; XGBoost; Extra Trees; stacking; feature selection; learning analytics.

How to cite this paper

Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning" Iconic Research And Engineering Journals Volume 10 Issue 3 2026 Page 534-537 https://doi.org/10.64388/IREV10I3-1722829
Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026, doi: https://doi.org/10.64388/IREV10I3-1722829
Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi (2026). Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning. Iconic Research And Engineering Journals, 10(3). doi: https://doi.org/10.64388/IREV10I3-1722829
Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi "Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning" Iconic Research And Engineering Journals, vol. 10, no. 3, Sep. 2026. Crossref, https://doi.org/10.64388/IREV10I3-1722829
@article{1722829,
      author = {Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi},
      title = {Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {3},
      pages = {534-537},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722829.pdf},
      abstract = {This study develops and empirically evaluates an ensemble machine-learning framework for dyslexia detection using Random-Forest Recursive Feature Elimination (RF-RFE), Random Forest (RF), XGBoost, Extra Trees (ET), and probability-level logistic-regression stacking. The analysis uses the supplied Dyt-desktop.csv dataset, which contains 3,644 observations, 196 predictors and one binary dyslexia outcome. The target distribution is imbalanced: 392 dyslexia cases (10.76%) and 3,252 non-dyslexia cases (89.24%). Five-fold stratified out-of-fold evaluation was used. Class-weighted RF, XGBoost and ET were compared with RF-RFE variants, while the proposed ensemble combined out-of-fold probability predictions from RF, XGBoost and ET through logistic regression. At the default 0.50 threshold, XGBoost produced the strongest individual accuracy (90.64%), while RF-RFE + XGBoost achieved 90.81% accuracy, 64.77% precision, 31.89% recall, 97.91% specificity, F1=0.4274, ROC-AUC=0.8858 and MCC=0.4122. The proposed stacking ensemble produced ROC-AUC=0.8839, recall=71.17%, F1=0.5171 and MCC=0.4644, but lower accuracy (85.70%) because its class-weighted meta-classifier generated substantially more positive predictions. These findings demonstrate that model ranking changes according to the evaluation objective: RF-RFE + XGBoost is strongest on accuracy and ROC-AUC, whereas the proposed ensemble provides the strongest minority-class recall, F1 and MCC at the tested threshold. The results support threshold optimization and precision-recall analysis as essential components of dyslexia screening.},
      keywords = {dyslexia; machine learning; ensemble learning; RF-RFE; Random Forest; XGBoost; Extra Trees; stacking; feature selection; learning analytics.},
      month = {September},
      doi = {https://doi.org/10.64388/IREV10I3-1722829}
  }