Home / Current Issue / Paper 1708157
Voice, Accent, And Identity in AI Interpreting: Toward More Inclusive Language Models
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence
Abstract
Artificial intelligence interpreting systems have gained significance because they remove language barriers by providing instant translation services, virtual help, and global communication channels. These systems experience difficulties detecting and providing a fair interpretation of various voice patterns and localized accents, contributing to automated system discrimination. This research evaluates the complex relationships between voices, accents, and identity schemes found in AI language models and the cultural and ethical risks of data training deficits. The research document stresses the necessity of inclusive voice-based AI systems because such measures protect against maintaining unfair stereotypes while fighting digital inequality. This research uses multiple social science approaches from computational linguistics with sociolinguistics and human-centered AI design to investigate current AI model interpretations of accents, the misrepresentation of identity, and its effects on speakers with non-standard dialects and minority languages. Current AI systems are analyzed using benchmarks, accent bias audit reports, and real-world interactions between AI models and humans. AI models demonstrate enduring precision differences based on accent variations, affecting exposure to AI tools and their reliability and equitable treatment in AI systems. Summary strategies appear to help develop AI interpreting systems that demonstrate cultural responsiveness and greater equity in their operation. AI system performance improvement requires three main actions: training dataset expansion, proper voice sampling standards, and team collaboration between technological developers and community representatives. A future language technology emerges through designing AI with voice and identity at its center, so it will respect all forms of language instead of eliminating them.
Keywords
Accent Bias in AI, Inclusive Language Models, Voice Recognition and Identity, Sociolinguistics in NLP, Ethical AI Interpreting
References
[1] Addressing the selection bias in voice assistance: training a voice assistance model in Python with equal data selection. (2023). Computer Vision Studies. https://doi.org/10.58396/cvs020103
[2] Alsaify, B. A., Arja, H. S. A., Maayah, B. Y., & Al-Taweel, M. M. (2022). A dataset for voice-based human identity recognition. Data in Brief, 42. https://doi.org/10.1016/j.dib.2022.108070
[3] Alshaabi, T., Dewhurst, D. R., Minot, J. R., Arnold, M. V., Adams, J. L., Danforth, C. M., & Dodds, P. S. (2021). The growing amplification of social media: measuring temporal and social contagion dynamics for over 150 languages on Twitter for 2009–2020. EPJ Data Science, 10(1). https://doi.org/10.1140/epjds/s13688-021-00271-0
[4] Borowiak, K., & von Kriegstein, K. (2020). Intranasal oxytocin modulates brain responses to voice-identity recognition in typically developing individuals, but not in ASD. Translational Psychiatry, 10(1). https://doi.org/10.1038/s41398-020-00903-5
[5] Bucholtz, M., & Hall, K. (2005). Identity and interaction: A sociocultural linguistic approach. Discourse Studies, 7(4–5), 585–614. https://doi.org/10.1177/1461445605054407
[6] Building a Pipeline of Bilingual SLPs to Serve Dual Language Learners: an Inclusive Model of Interprofessional Education. (2021). Teaching and Learning in Communication Sciences and Disorders. https://doi.org/10.30707/tlcsd5.3.1649037688.673503
[7] Cong, Y. (2022). Pre-trained Language Models’ Interpretation of Evaluativity Implicature: Evidence from Gradable Adjectives Usage in Context. In UnImplicit 2022 - 2nd Workshop on Understanding Implicit and Underspecified Language, Proceedings of the Workshop (pp. 1–7). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.unimplicit-1.1
[8] Couper-Kuhlen, E. (2004). Prosody and sequence organization in English conversation. In E. Couper-Kuhlen & C. E. Ford (Eds.), Sound patterns in interaction: Cross-linguistic studies from conversation (pp. 335–376). John Benjamins. https://doi.org/10.1075/sidag.1
[9] David, R. D., & Anderson, C. E. (2022). The Universal Genre Sphere: A Curricular Model Integrating GBA and UDL to Promote Equitable Academic Writing Instruction for EAL University Students. Center for Educational Policy Studies Journal, 12(4), 35–52. https://doi.org/10.26529/cepsj.1442
[10] ewis, A. (2022). Multimodal large language models for inclusive collaboration learning tasks. In NAACL 2022 - 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Student Research Workshop (pp. 202–210). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.naacl-srw.26
[11] Haut, K., Wohn, C., Antony, V., Goldfarb, A., Welsh, M., Sumanthiran, D., … Hoque, E. (2022). Demographic Feature Isolation for Bias Research using Deepfakes. In MM 2022 - Proceedings of the 30th ACM International Conference on Multimedia (pp. 6890–6897). Association for Computing Machinery, Inc. https://doi.org/10.1145/3503161.3549204
[12] Horváth, I. (2022). AI in interpreting: Ethical considerations. Across Languages and Cultures, 23(1), 1–13. https://doi.org/10.1556/084.2022.00108
[13] Huang, Z., Wu, Y., Zhang, Z., Zhou, H., & Xie, L. (2021). Cross-lingual accented speech recognition with adversarial training and data augmentation. IEEE Transactions on Audio, Speech, and Language Processing, 29, 2346–2357.https://doi.org/10.1109/TASLP.2021.308916
[14] Jandrić, P. (2019). The Postdigital Challenge of Critical Media Literacy. The International Journal of Critical Media Literacy, 1(1), 26–37. https://doi.org/10.1163/25900110-00101002
[15] Jia, Y., Zhang, Y., Weiss, R. J., Wang, Q., Shen, J., Ren, F., ... & Wu, Y. (2018). Transfer learning from speaker verification to multispeaker text-to-speech synthesis. arXiv preprint. https://doi.org/10.48550/arXiv.1806.04558
[16] Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., ... & Goel, S. (2020). Racial disparities in automated speech recognition. Proceedings of the National Academy of Sciences, 117(14), 7684–7689.https://doi.org/10.1073/pnas.1915768117
[17] Labov, W. (1972). Sociolinguistic patterns. University of Pennsylvania Press.https://doi.org/10.9783/9780812200909
[18] Lippi-Green, R. (1997). English with an accent: Language, ideology, and discrimination in the United States. Routledge. https://doi.org/10.4324/9781315844393
[19] Maguinness, C., & von Kriegstein, K. (2021). Visual mechanisms for voice-identity recognition flexibly adjust to auditory noise level. Human Brain Mapping, 42(12), 3963–3982. https://doi.org/10.1002/hbm.25532
[20] Manggiasih, L. A., Loreana, Y. R., Azizah, A., & Nurjati, N. (2023). Strengths and Limitations of the SmallTalk2Me App in English Language Proficiency Evaluation. Tell : Teaching of English Language and Literature Journal, 11(2). https://doi.org/10.30651/tell.v11i2.19560
[21] Nguyen, D., Rosseel, L., & Grieve, J. (2021). On learning and representing social meaning in NLP: a sociolinguistic perspective. In NAACL-HLT 2021 - 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference (pp. 603–612). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2021.naacl-main.50
[22] Qiao-Franco, G., & Zhu, R. (2024). China’s Artificial Intelligence Ethics: Policy Development in an Emergent Community of Practice. Journal of Contemporary China, 33(146), 189–205. https://doi.org/10.1080/10670564.2022.2153016
[23] Roswandowitz, C., Kappes, C., Obrig, H., & Von Kriegstein, K. (2018). Obligatory and facultative brain regions for voice-identity recognition. Brain, 141(1), 234–247. https://doi.org/10.1093/brain/awx313
[24] Snoddon, K., & Murray, J. J. (2023). Supporting deaf learners in Nepal via Sustainable Development Goal 4: Inclusive and equitable quality education in sign languages. International Journal of Speech-Language Pathology. Taylor and Francis Ltd. https://doi.org/10.1080/17549507.2022.2141325
[25] Tatman, R. (2017). Gender and dialect bias in YouTube’s automatic captions. Proceedings of the First ACL Workshop on Ethics in Natural Language Processing, 53–59. https://doi.org/10.18653/v1/W17-1606
How to cite this paper
@article{1708157,
author = {Dilshat Azizov},
title = {Voice, Accent, And Identity in AI Interpreting: Toward More Inclusive Language Models},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {6},
pages = {498-506},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1708157.pdf},
abstract = {Artificial intelligence interpreting systems have gained significance because they remove language barriers by providing instant translation services, virtual help, and global communication channels. These systems experience difficulties detecting and providing a fair interpretation of various voice patterns and localized accents, contributing to automated system discrimination. This research evaluates the complex relationships between voices, accents, and identity schemes found in AI language models and the cultural and ethical risks of data training deficits. The research document stresses the necessity of inclusive voice-based AI systems because such measures protect against maintaining unfair stereotypes while fighting digital inequality. This research uses multiple social science approaches from computational linguistics with sociolinguistics and human-centered AI design to investigate current AI model interpretations of accents, the misrepresentation of identity, and its effects on speakers with non-standard dialects and minority languages. Current AI systems are analyzed using benchmarks, accent bias audit reports, and real-world interactions between AI models and humans. AI models demonstrate enduring precision differences based on accent variations, affecting exposure to AI tools and their reliability and equitable treatment in AI systems. Summary strategies appear to help develop AI interpreting systems that demonstrate cultural responsiveness and greater equity in their operation. AI system performance improvement requires three main actions: training dataset expansion, proper voice sampling standards, and team collaboration between technological developers and community representatives. A future language technology emerges through designing AI with voice and identity at its center, so it will respect all forms of language instead of eliminating them.},
keywords = {Accent Bias in AI, Inclusive Language Models, Voice Recognition and Identity, Sociolinguistics in NLP, Ethical AI Interpreting},
month = {December},
}