Home / Current Issue / Paper 1722829
Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning
Subject area: Science,Engineering and Technology · Area of research: Machine Learning
DOI: https://doi.org/10.64388/IREV10I3-1722829
Abstract
This study develops and empirically evaluates an ensemble machine-learning framework for dyslexia detection using Random-Forest Recursive Feature Elimination (RF-RFE), Random Forest (RF), XGBoost, Extra Trees (ET), and probability-level logistic-regression stacking. The analysis uses the supplied Dyt-desktop.csv dataset, which contains 3,644 observations, 196 predictors and one binary dyslexia outcome. The target distribution is imbalanced: 392 dyslexia cases (10.76%) and 3,252 non-dyslexia cases (89.24%). Five-fold stratified out-of-fold evaluation was used. Class-weighted RF, XGBoost and ET were compared with RF-RFE variants, while the proposed ensemble combined out-of-fold probability predictions from RF, XGBoost and ET through logistic regression. At the default 0.50 threshold, XGBoost produced the strongest individual accuracy (90.64%), while RF-RFE + XGBoost achieved 90.81% accuracy, 64.77% precision, 31.89% recall, 97.91% specificity, F1=0.4274, ROC-AUC=0.8858 and MCC=0.4122. The proposed stacking ensemble produced ROC-AUC=0.8839, recall=71.17%, F1=0.5171 and MCC=0.4644, but lower accuracy (85.70%) because its class-weighted meta-classifier generated substantially more positive predictions. These findings demonstrate that model ranking changes according to the evaluation objective: RF-RFE + XGBoost is strongest on accuracy and ROC-AUC, whereas the proposed ensemble provides the strongest minority-class recall, F1 and MCC at the tested threshold. The results support threshold optimization and precision-recall analysis as essential components of dyslexia screening.
Keywords
dyslexia; machine learning; ensemble learning; RF-RFE; Random Forest; XGBoost; Extra Trees; stacking; feature selection; learning analytics.
How to cite this paper
@article{1722829,
author = {Patrick Ufomba Nwogu, Prof. Chigozirim Ajaegu, Dr. Faruk Umar Ambursa, Dr. Femi Adeluyi},
title = {Exploratory Analysis of Ensemble Machine Learning for Dyslexia Detection Using RF-RFE and Stacked Learning},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {10},
number = {3},
pages = {534-537},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722829.pdf},
abstract = {This study develops and empirically evaluates an ensemble machine-learning framework for dyslexia detection using Random-Forest Recursive Feature Elimination (RF-RFE), Random Forest (RF), XGBoost, Extra Trees (ET), and probability-level logistic-regression stacking. The analysis uses the supplied Dyt-desktop.csv dataset, which contains 3,644 observations, 196 predictors and one binary dyslexia outcome. The target distribution is imbalanced: 392 dyslexia cases (10.76%) and 3,252 non-dyslexia cases (89.24%). Five-fold stratified out-of-fold evaluation was used. Class-weighted RF, XGBoost and ET were compared with RF-RFE variants, while the proposed ensemble combined out-of-fold probability predictions from RF, XGBoost and ET through logistic regression. At the default 0.50 threshold, XGBoost produced the strongest individual accuracy (90.64%), while RF-RFE + XGBoost achieved 90.81% accuracy, 64.77% precision, 31.89% recall, 97.91% specificity, F1=0.4274, ROC-AUC=0.8858 and MCC=0.4122. The proposed stacking ensemble produced ROC-AUC=0.8839, recall=71.17%, F1=0.5171 and MCC=0.4644, but lower accuracy (85.70%) because its class-weighted meta-classifier generated substantially more positive predictions. These findings demonstrate that model ranking changes according to the evaluation objective: RF-RFE + XGBoost is strongest on accuracy and ROC-AUC, whereas the proposed ensemble provides the strongest minority-class recall, F1 and MCC at the tested threshold. The results support threshold optimization and precision-recall analysis as essential components of dyslexia screening.},
keywords = {dyslexia; machine learning; ensemble learning; RF-RFE; Random Forest; XGBoost; Extra Trees; stacking; feature selection; learning analytics.},
month = {September},
doi = {https://doi.org/10.64388/IREV10I3-1722829}
}