International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1714681

1714681PublishedVol 9 · Issue 8

Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning.

Muhammad Awwal Ibrahim DT Chinyio

Subject area: Science,Engineering and Technology  ·  Area of research: Computer Science

DOI: https://doi.org/10.64388/IREV9I8-1714681

Abstract

The rise of social media has intensified cyberbullying, impacting users across diverse languages, yet low-resource languages like Hausa lack effective detection tools. With over 100 million Hausa speakers predominantly in Nigeria and Niger, addressing this gap is crucial for fostering safer online environments. This study aims to develop a machine learning model to detect cyberbullying in Hausa tweets on X (formerly Twitter). The methodology involved collecting and annotating Hausa-language tweets, preprocessing the data through cleaning and removal of stopwords using Natural Language Tokenization (NLTK) Library, and extracting features using TF-IDF technique. Multiple classifiers, including Support Vector Machine (SVM), XGBoost, and Logistic Regression, were trained and evaluated based on accuracy, precision, and recall. Results showed that the SVM outperformed others with an accuracy of 0.98, followed by Random Forest 0.9794, then XGBoost 0.9769, while Logistic Regression had the lowest accuracy 0.9527. The findings demonstrate that culturally-sensitive, language-specific models can enhance cyberbullying detection in Hausa, contributing to safer digital spaces and providing a basis for further research in low-resource language NLP applications.

Keywords

Support Vector Machine, Tweets on X XGBoost, Natural Language Tokenization

How to cite this paper

Muhammad Awwal Ibrahim, DT Chinyio "Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning." Iconic Research And Engineering Journals Volume 9 Issue 8 2026 Page 2018-2032 https://doi.org/10.64388/IREV9I8-1714681
Muhammad Awwal Ibrahim, DT Chinyio "Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning." Iconic Research And Engineering Journals, vol. 9, no. 8, Feb. 2026, doi: https://doi.org/10.64388/IREV9I8-1714681
Muhammad Awwal Ibrahim, DT Chinyio (2026). Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning.. Iconic Research And Engineering Journals, 9(8). doi: https://doi.org/10.64388/IREV9I8-1714681
Muhammad Awwal Ibrahim, DT Chinyio "Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning." Iconic Research And Engineering Journals, vol. 9, no. 8, Feb. 2026. Crossref, https://doi.org/10.64388/IREV9I8-1714681
@article{1714681,
      author = {Muhammad Awwal Ibrahim, DT Chinyio},
      title = {Cyberbullying Detection in Hausa Language on X Social Medium Using Machine Learning.},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {8},
      pages = {2018-2032},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1714681.pdf},
      abstract = {The rise of social media has intensified cyberbullying, impacting users across diverse languages, yet low-resource languages like Hausa lack effective detection tools. With over 100 million Hausa speakers predominantly in Nigeria and Niger, addressing this gap is crucial for fostering safer online environments. This study aims to develop a machine learning model to detect cyberbullying in Hausa tweets on X (formerly Twitter). The methodology involved collecting and annotating Hausa-language tweets, preprocessing the data through cleaning and removal of stopwords using Natural Language Tokenization (NLTK) Library, and extracting features using TF-IDF technique. Multiple classifiers, including Support Vector Machine (SVM), XGBoost, and Logistic Regression, were trained and evaluated based on accuracy, precision, and recall. Results showed that the SVM outperformed others with an accuracy of 0.98, followed by Random Forest 0.9794, then XGBoost 0.9769, while Logistic Regression had the lowest accuracy 0.9527. The findings demonstrate that culturally-sensitive, language-specific models can enhance cyberbullying detection in Hausa, contributing to safer digital spaces and providing a basis for further research in low-resource language NLP applications.},
      keywords = {Support Vector Machine, Tweets on X XGBoost, Natural Language Tokenization },
      month = {February},
      doi = {https://doi.org/10.64388/IREV9I8-1714681}
  }