Home / Current Issue / Paper 1718336
HealStation AI: A Multimodal Medical Consultation System Integrating Large Language Models, Voice Processing, and Location-Based Healthcare Discovery
Subject area: Biological & Medical Sciences · Area of research: Artificial Intelligence in Healthcare
DOI: https://doi.org/10.64388/IREV9I11-1718336
Abstract
Access to timely, qualified medical consultation remains inequitably distributed across socioeconomic and geographic boundaries. In India, the physician-to-population ratio falls well below the World Health Organization’s recommended threshold, creating critical delays in triage and specialist referral. This paper presents HealStation AI, an open-source multimodal medical consultation platform engineered to bridge this gap by combining state-of-the-art large language models (LLMs), automatic speech recognition (ASR), text-to-speech synthesis, document vision analysis, and real-time location-based healthcare provider discovery. The system integrates Meta’s Llama-4-Scout-17B multimodal model and Llama-3.3-70B via the Groq inference API, Microsoft Edge-TTS for naturalistic voice synthesis, and OpenStreetMap-powered facility lookup through LocationIQ to deliver end-to-end patient consultation workflows. HealStation AI supports three Indian languages—English, Hindi, and Marathi—processes medical images and PDF laboratory reports, performs rule-based emergency triage using validated clinical scoring instruments (HEART score, BE-FAST, trauma scoring), and recommends appropriate specialists from a taxonomy of twenty clinical domains. A FastAPI backend exposes a well-defined REST API consumed by a React 19 single-page application. Empirical test cases across low, medium, and high urgency symptom profiles demonstrate consistent specialist routing and urgency classification with an end-to-end latency of 6.2 seconds (N=50). The platform is designed for deployment in resource-constrained environments and is structured for extensibility toward telemedicine and Ayushman Bharat Digital Mission (ABDM) integration.
Keywords
Medical Artificial Intelligence, Large Language Models, Multimodal Health Systems, Automatic Speech Recognition, Location-Based Healthcare, Emergency Triage, Multilingual NLP, Telemedicine, Fastapi, React.
References
[1] World Health Organization, “Health workforce: Doctor density per 10,000 population,” Global Health Observatory Data Repository, 2023. [Online]. Available: https://www.who.int/data/gho
[2] K. Singhal, S. Azizi, T. Tu, et al., “Large language models encode clinical knowledge,” Nature, vol. 620, pp. 172–180, 2023. doi:10.1038/s41586-023-06291-2
[3] K. Singhal, T. Tu, J. Gottweis, et al., “Towards expert-level medical question answering with large language models,” arXiv:2305.09617, 2023.
[4] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML, 2023.
[5] OpenAI, “GPT-4 Technical Report,” arXiv:2303.08774, 2023.
[6] C. Li, C. Wong, S. Zhang, et al., “LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day,” in Proc. NeurIPS Datasets and Benchmarks Track, 2023.
[7] K. Bannur, S. Hyland, Q. Liu, et al., “Learning to exploit temporal structure for biomedical vision-language processing,” in Proc. CVPR, 2023.
[8] M. Guagliardo, “Theoretical grounds and measurement methods of geographic accessibility to health services,” SESPAS Gac Sanit, vol. 19, no. 4, pp. 1–18, 2004.
[9] Meta AI, “Llama 4: The next generation of open large language models,” Meta Blog, 2025. [Online]. Available: https://ai.meta.com/llama
[10] Groq Inc., “GroqCloud: LPU inference engine,” Technical Documentation, 2024. [Online]. Available: https://console.groq.com/docs
[11] J. Lee, W. Yoon, S. Kim, et al., “BioBERT: A pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics, vol. 36, no. 4, pp. 1234–1240, 2020.
[12] LocationIQ, “Places API: Medical facilities search,” API Documentation, 2024. [Online]. Available: https://locationiq.com/docs
[13] OpenStreetMap Foundation, “OpenStreetMap Wiki: Medical facilities tagging schema,” 2024. [Online]. Available: https://wiki.openstreetmap.org/wiki/Healthcare
[14] V. Shah, V. Bhatt, S. Singh, and R. Kumar, “AI-based health chatbots and their role in the Indian healthcare ecosystem: A systematic review,” Health Informatics J., vol. 28, no. 3, 2022. doi:10.1177/14604582221112321
[15] T. B. Brown, B. Mann, N. Ryder, et al., “Language models are few-shot learners,” in Proc. NeurIPS, vol. 33, pp. 1877–1901, 2020.
[16] S. Sebastian and B. Balachandran, “Whisper-Hindi: Domain adaptation of automatic speech recognition for Hindi medical transcription,” arXiv:2311.10240, 2023.
[17] G. Kaissis, M. Makowski, D. Rückert, and R. Braren, “Secure, privacy-preserving and federated machine learning in medical imaging,” Nat. Mach. Intell., vol. 2, pp. 305–311, 2020.
[18] Ministry of Health and Family Welfare, India, “Ayushman Bharat Digital Mission: Technical specifications,” NHA Report, 2023. [Online]. Available: https://abdm.gov.in
[19] S. Tipirneni and C. K. Reddy, “Self-supervised transformer for sparse and irregularly sampled multivariate clinical time-series,” ACM Trans. Knowl. Discov. Data, vol. 16, no. 4, pp. 1–17, 2022.
[20] P. Rajpurkar, J. Irvin, K. Ball, et al., “CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning,” arXiv:1711.05225, 2017.
How to cite this paper
@article{1718336,
author = {Abhijeet Joshi, Dr. P. D. Adkar},
title = {HealStation AI: A Multimodal Medical Consultation System Integrating Large Language Models, Voice Processing, and Location-Based Healthcare Discovery},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {4667-4674},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718336.pdf},
abstract = {Access to timely, qualified medical consultation remains inequitably distributed across socioeconomic and geographic boundaries. In India, the physician-to-population ratio falls well below the World Health Organization’s recommended threshold, creating critical delays in triage and specialist referral. This paper presents HealStation AI, an open-source multimodal medical consultation platform engineered to bridge this gap by combining state-of-the-art large language models (LLMs), automatic speech recognition (ASR), text-to-speech synthesis, document vision analysis, and real-time location-based healthcare provider discovery. The system integrates Meta’s Llama-4-Scout-17B multimodal model and Llama-3.3-70B via the Groq inference API, Microsoft Edge-TTS for naturalistic voice synthesis, and OpenStreetMap-powered facility lookup through LocationIQ to deliver end-to-end patient consultation workflows. HealStation AI supports three Indian languages—English, Hindi, and Marathi—processes medical images and PDF laboratory reports, performs rule-based emergency triage using validated clinical scoring instruments (HEART score, BE-FAST, trauma scoring), and recommends appropriate specialists from a taxonomy of twenty clinical domains. A FastAPI backend exposes a well-defined REST API consumed by a React 19 single-page application. Empirical test cases across low, medium, and high urgency symptom profiles demonstrate consistent specialist routing and urgency classification with an end-to-end latency of 6.2 seconds (N=50). The platform is designed for deployment in resource-constrained environments and is structured for extensibility toward telemedicine and Ayushman Bharat Digital Mission (ABDM) integration.},
keywords = {Medical Artificial Intelligence, Large Language Models, Multimodal Health Systems, Automatic Speech Recognition, Location-Based Healthcare, Emergency Triage, Multilingual NLP, Telemedicine, Fastapi, React.},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1718336}
}