International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1713136

1713136 Vol 9 · Issue 6 Download Paper

The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion

Nneoma Udeze

Subject area: Arts, Social Sciences and Humanities  ·  Area of research: Digital Inclusion and Language Preservation

DOI: 10.64388/IREV9I6-1713136

Abstract

The computer era has led to new opportunities in communication and knowledge sharing that were never had before, and it has also endangered the diversity of languages in the world and increased the disparity in technology. As more languages in the world are at risk of extinction (estimated to be 40 percent), and many communities do not meaningfully access digital tools in their native languages, the importance of machine learning in language preservation is greater. This study discussed the application of machine learning technologies to document, support, and revive endangered languages and ensure digital inclusion. Automatic speech recognition had a major effect of reducing the time spent in transcription of endangered languages to almost an instant output with just a few training samples. Machine translation systems, such as new multilingual projects, have increased their languages to cover more than 200 languages and enhanced the quality of translations done on languages that were previously ignored. In addition to documentation, digital platforms, educational technologies, and available language resources that helped minority language communities were also developed because of machine learning. Overall, this paper demonstrated that machine learning offers useful means to solve linguistic vulnerability and cut digital inequality in the ever-connected world.

Keywords

Machine Learning, Language Preservation, Digital Inclusion, Endangered Languages, Natural Language Processing, Language Diversity, Old Data Sovereignty, Limited Resource Languages, And Digital Equity.

References

[1] Adams, O., Cohn, T., Neubig, G., Cruz, H., Bird, S., & Michaud, A. (2018). Evaluating phonemic transcription of low-resource tonal languages for language documentation. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (pp. 3356–3365). European Language Resources Association. http://www.lrec-conf.org/proceedings/lrec2018/pdf/344.pdf

[2] Anderson, G. (2011). Language hotspots: What (applied) linguistics and education should do about language endangerment in the twenty-first century. Language and Education, 25(4), 273–289. https://doi.org/10.1080/09500782.2011.577218

[3] Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587–604. https://doi.org/10.1162/tacl_a_00041

[4] Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.463

[5] Blasi, D., Anastasopoulos, A., & Neubig, G. (2022). Systematic inequalities in language technology performance across the world’s languages. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5486–5505. https://doi.org/10.18653/v1/2022.acl-long.376

[6] Blodgett, S. L., Barocas, S., Daumé III, H., & Wallach, H. (2020). Language (technology) is power: A critical survey of “bias” in NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5454–5476. https://doi.org/10.18653/v1/2020.acl-main.485

[7] Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, É., Ott, M., Zettlemoyer, L., & Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 8440–8451. https://doi.org/10.18653/v1/2020.acl-main.747

[8] Facer, K., & Selwyn, N. (2021). Digital technology and the futures of education: Towards ‘non-stupid’ optimism (ED-2020/FoE-BP/27). UNESCO.

[9] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. http://www.deeplearningbook.org

[10] Himmelmann, N. (2006). Language documentation: What it is and what it is good for. In J. Gippert, N. Himmelmann, & U. Mosel (Eds.), Essentials of language documentation (pp. 1–30). Mouton de Gruyter.

[11] Holmes, W., Persson, J., Chounta, I.-A., Wasson, B., & Dimitrova, V. (2022). Artificial intelligence and education: A critical view through the lens of human rights, democracy and the rule of law. Council of Europe Publishing. https://rm.coe.int/1680a956e3

[12] Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The state and fate of linguistic diversity and inclusion in the NLP world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 6282–6293). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.560

[13] Khan, W. N. (2024.). Ethical challenges of AI in education: Balancing innovation with data privacy (pp. 1–13) [Unpublished manuscript]. NCBA.

[14] Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689. https://doi.org/10.1073/pnas.1915768117

[15] Leonard, W. (2017). Producing language reclamation by decolonising “language”. Language Documentation and Description, 14. https://doi.org/10.25894/ldd146

[16] Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2). https://doi.org/10.1177/2053951716679679

[17] Moseley, C. (Ed.). (2010). Atlas of the world's languages in danger (3rd ed.). UNESCO Publishing. https://uploads.guim.co.uk/2024/09/30/UNESCO_Atlas_of_Languages_2010.pdf

[18] Orife, I., Kreutzer, J., Sibanda, B., Whitenack, D., Siminyu, K., Martinus, L., Ali, J. T., Abbott, J., Marivate, V., Kabongo, S., Meressa, M., Murhabazi, E., Ahia, O., van Biljon, E., Ramkilowan, A., Akinfaderin, A., Öktem, A., Akin, W., Kioko, G., Degila, K., . . . Bashir, A. (2020). Masakhane -- Machine translation for Africa [Conference paper]. AfricaNLP Workshop, ICLR 2020. https://doi.org/10.48550/arXiv.2003.11529

[19] Pedro, F., Subosa, M., Rivas, A., & Valverde, P. (2019). Artificial intelligence in education: Challenges and opportunities for sustainable development. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000366994

[20] Piller, I. (2016). Linguistic diversity and social justice: An introduction to applied sociolinguistics. Oxford University Press. https://doi.org/10.1093/acprof:oso/9780199937240.001.0001

[21] Roll, I., & Wylie, R. (2016). Evolution and revolution in artificial intelligence in education. International Journal of Artificial Intelligence in Education, 26, 582–599. https://doi.org/10.1007/s40593-016-0110-3

[22] Ruíz, R. (1984). Orientations in language planning. NABE Journal, 8(2), 15–34. https://doi.org/10.1080/08855072.1984.10668464

[23] Siemens, G. (2005). Connectivism: A learning theory for the digital age. International Journal of Instructional Technology and Distance Learning, 2(1). http://www.itdl.org/Journal/Jan_05/article01.htm

[24] Simons, G. F., & Fennig, C. D. (Eds.). (2018). Ethnologue: Languages of the world (21st ed.). SIL International. http://www.ethnologue.com

[25] Tojimuxammadov, J. (2025). Ethical challenges of artificial intelligence in education. Scientia: Technology, Science and Society, 2, 90–96. https://doi.org/10.59324/stss.2025.2(11).09

How to cite this paper

Nneoma Udeze "The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion" Iconic Research And Engineering Journals Volume 9 Issue 6 2025 Page 2335-2340 https://doi.org/10.64388/IREV9I6-1713136
Nneoma Udeze "The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion" Iconic Research And Engineering Journals, vol. 9, no. 6, Dec. 2025, doi: https://doi.org/10.64388/IREV9I6-1713136
Nneoma Udeze (2025). The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion. Iconic Research And Engineering Journals, 9(6). doi: https://doi.org/10.64388/IREV9I6-1713136
Nneoma Udeze "The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion" Iconic Research And Engineering Journals, vol. 9, no. 6, Dec. 2025. Crossref, https://doi.org/10.64388/IREV9I6-1713136
@article{1713136,
      author = {Nneoma Udeze},
      title = {The Role of Machine Learning in Preserving Languages and Promoting Digital Inclusion},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {6},
      pages = {2335-2340},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1713136.pdf},
      abstract = {The computer era has led to new opportunities in communication and knowledge sharing that were never had before, and it has also endangered the diversity of languages in the world and increased the disparity in technology. As more languages in the world are at risk of extinction (estimated to be 40 percent), and many communities do not meaningfully access digital tools in their native languages, the importance of machine learning in language preservation is greater. This study discussed the application of machine learning technologies to document, support, and revive endangered languages and ensure digital inclusion. Automatic speech recognition had a major effect of reducing the time spent in transcription of endangered languages to almost an instant output with just a few training samples. Machine translation systems, such as new multilingual projects, have increased their languages to cover more than 200 languages and enhanced the quality of translations done on languages that were previously ignored. In addition to documentation, digital platforms, educational technologies, and available language resources that helped minority language communities were also developed because of machine learning. Overall, this paper demonstrated that machine learning offers useful means to solve linguistic vulnerability and cut digital inequality in the ever-connected world.},
      keywords = {Machine Learning, Language Preservation, Digital Inclusion, Endangered Languages, Natural Language Processing, Language Diversity, Old Data Sovereignty, Limited Resource Languages, And Digital Equity.},
      month = {December},
      doi = {https://doi.org/10.64388/IREV9I6-1713136}
  }