International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1710732

1710732 Vol 9 · Issue 3 Download Paper

Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings

Dr. P. A. Manoj Kumar Dileep Parasu Sarvesh Shashikumar Kalyan Guru

Subject area: Science,Engineering and Technology  ·  Area of research: Prompt , AI in Medicine

DOI: 10.64388/IREV9I3-1710732-2258

Abstract

Background: Indian emergency medicine is now under severe shortage of physicians, face a surge of patients and broad language barriers which require new AI methods that respond in the local health environments. Study methods: In the study, we implemented a multi-center randomized controlled trial in several leading academic medical centers in India between 01 January 2023 and 31 December 2023. We did the study on 1,000 adult emergency department patients and assigned them to AI-assisted care (ChatGPT-4 with culturally-adapted prompt engineering) vs. standard care. Length of stay, adverse events at 30 days, and diagnostics accuracy were used as primary endpoints. The quality of AI-produced clinical summaries was measured with ROUGE, BLEU, LSA metrics and compared to the way it was documented by the physician. There was blindness to treatment in all the outcome assessors. Findings: Of 1,000 randomized individuals (500 AI-assisted, 500 standard care) non-inferiority of AI-assisted care was shown in diagnostic accuracy (AI-assisted care 94.8%; standard care 94.2%; difference 0.6%, 95% CI: -2.1 to 3.3), and AI-assisted care had a superior performance in length of stay (AI-assisted care median 3.1; standard care median 4.3 hours; difference -1.2 hours). The Natural language processing evaluation showed high agreement, ROUGE-L scores 0.862?0.11, ROUGE-2F scores 0.804?0.14, and 689/1,000 (68.9%) cases scored 0.85 or above. The use of AI-assisted care saved physicians 38 percent of documentation time (P<0.001), raised clinical guidelines compliance by 23 percent (P<0.001) and raised patient satisfaction ratings (8.6 vs. 7.8; P<0.001). The cost-effectiveness analysis displayed savings of 2,847 Indian rupees per patient. Conclusions: It was accomplished with excellent clinical results and outstanding financial cost savings, relative to the clinical outcomes, cultural adaptations of prompt engineering and AI-aided emergency care. These results confirm the use of AI nationwide in Indian emergency medicine.

Keywords

Artificial Intelligence, Emergency Medicine, Prompt Engineering, Natural Language Processing, ROUGE Score, Clinical Decision Support, Indian Healthcare, Cultural Adaptation

References

[1] Barnes, A. J., Zhang, Y., & Valenzuela, A. (2024). AI and culture: Culturally dependent responses to AI systems. Current Opinion in Psychology, 58, 101838. https://doi.org/10.1016/j.copsyc.2024.101838

[2] Bazzano, A., Mantsios, A., Mattei, N., Kosorok, M., & Culotta, A. (2025). AI can be a powerful social innovation for public health if community engagement is at the core. Journal of Medical Internet Research, 27, e68198. https://doi.org/10.2196/68198

[3] Beam, A. L., & Kohane, I. S. (2018). Big data and machine learning in health care. JAMA, 319(13), 1317-1318.

[4] Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623). https://doi.org/10.1145/3442188.3445922

[5] Bharadwaj, P., Nicola, L., Breau-Brunel, M., Sensini, F., Tanova-Yotova, N., Atanasov, P., Lobig, F., & Blankenburg, M. (2024). Unlocking the value: Quantifying the return on investment of hospital artificial intelligence. Journal of the American College of Radiology, 21(10), 1677-1685.

[6] Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.

[7] Chada, B. V., & Summers, L. (2022). AI in the NHS: a framework for adoption. Future healthcare journal, 9(3), 313–316. https://doi.org/10.7861/fhj.2022-0068

[8] Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in healthcare. Annual Review of Biomedical Data Science, 4, 123-144. https://doi.org/10.1146/annurev-biodatasci-092820-114757

[9] Chouten, B. C., Cox, A., Duran, G., Kerremans, K., Banning, L. K., Lahdidioui, A., van den Muijsenbergh, M., Schinkel, S., Sungur, H., Suurmond, J., Zendedel, R., & Krystallidou, D. (2020). Mitigating language and cultural barriers in healthcare communication: Toward a holistic approach. Patient education and counseling, S0738-3991(20)30242-1. Advance online publication.

[10] Citarella, A. A., Barbella, M., Ciobanu, M. G., De Marco, F., Di Biasi, L., & Tortora, G. (2025). Assessing the effectiveness of ROUGE as unbiased metric in extractive vs. abstractive summarization techniques. Journal of Computational Science, 87, 102571. https://doi.org/10.1016/j.jocs.2025.102571

[11] Davenport, T., & Kalakota, R. (2019). The potential for artificial intelligence in healthcare. Future Healthcare Journal, 6(2), 94-98. https://doi.org/10.7861/futurehosp.6-2-94

[12] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171-4186).

[13] Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24-29. https://doi.org/10.1038/s41591-018-0316-z

[14] Fernandes, M., Vieira, S. M., Leite, F., Palos, C., Johnson, A., Finkelstein, S., Sirgo, G., Sousa, J. M. C., & Makse, H. A. (2020). Clinical decision support systems for triage in the emergency department using intelligent systems: A review. Artificial Intelligence in Medicine, 102, 101762.

[15] Goldberg, Y. (2016). A primer on neural network models for natural language processing. Journal of Artificial Intelligence Research, 57(1), 345–420.

[16] Gomez-Cabello, C. A., Borna, S., Pressman, S., Haider, S. A., Haider, C. R., & Forte, A. J. (2024). Artificial-Intelligence-Based Clinical Decision Support Systems in Primary Care: A Scoping Review of Current Clinical Implementations. European journal of investigation in health, psychology and education, 14(3), 685–698.

[17] Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., Venugopalan, S., Widner, K., Madams, T., Cuadros, J., Kim, R., Raman, R., Nelson, P. C., Mega, J. L., & Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA, 316(22), 2402-2410. https://doi.org/10.1001/jama.2016.17216

[18] Health Technology Assessment in India. (2025). Health technology assessment of AI-assisted CXR for interpretation for tuberculosis: A rapid health technology assessment . Indian Institute of Public Health Gandhinagar.

[19] Hirosawa, T., Harada, Y., Yokose, M., Sakamoto, T., Kawamura, R., & Shimizu, T. (2023). Diagnostic accuracy of differential-diagnosis lists generated by generative pretrained transformer 3 chatbot for clinical vignettes with common chief complaints: A pilot study. International Journal of Environmental Research and Public Health, 20(4), 3378.

[20] Hovy, D., & Spruit, S. L. (2016). The social impact of natural language processing. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 591-598). https://doi.org/10.18653/v1/P16-2096

[21] Jiang, F., Jiang, Y., Zhi, H., Dong, Y., Li, H., Ma, S., Wang, Y., Dong, Q., Shen, H., & Wang, Y. (2017). Artificial intelligence in healthcare: Past, present and future. Stroke and Vascular Neurology, 2(4), 230-243. https://doi.org/10.1136/svn-2017-000101

[22] Jones, K. S. (2007). Automatic summarising: The state of the art. Information Processing & Management, 43(6), 1449-1481. https://doi.org/10.1016/j.ipm.2007.03.009

[23] Kaczmarczyk, R., Wilhelm, T.I., Martin, R. et al. Evaluating multimodal AI in medical diagnostics. npj Digit. Med. 7, 205 (2024). https://doi.org/10.1038/s41746-024-01208-3

[24] Kakatum Rao, S., Gupta, P., Mohammed, A., Zakhmi, K., Ranjan Mohanty, M., & Prasad Jalaja, P. (2025). The Impact of Artificial Intelligence on Financial Systems in Healthcare: A Systematic Review of Economic Evaluation Studies. Cureus, 17(6), e86279. https://doi.org/10.7759/cureus.86279

[25] Lin, C. Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out: Proceedings of the ACL-04 Workshop (pp. 74-81). Association for Computational Linguistics.

[26] Madani, A., Arnaout, R., Mofrad, M., & Arnaout, R. (2018). Fast and accurate view classification of echocardiograms using deep learning. NPJ Digital Medicine, 1, 6. https://doi.org/10.1038/s41746-017-0013-1

[27] Miller, R. A. (1994). Medical diagnostic decision support systems—past, present, and future: A threaded bibliography and brief commentary. Journal of the American Medical Informatics Association, 1(1), 8-27. https://doi.org/10.1136/jamia.1994.95236141

[28] Mishra, R., & Shridevi, S. (2024). Knowledge graph driven medicine recommendation system using graph neural networks on longitudinal medical records. Scientific reports, 14(1), 25449. https://doi.org/10.1038/s41598-024-75784-5

[29] Naderbagi, A., Loblay, V., Zahed, I., Ekambareshwar, M., Poulsen, A., Song, Y., Ospina-Pinillos, L., Krausz, M., Mamdouh Kamel, M., Hickie, I., & LaMonica, H. (2024). Cultural and contextual adaptation of digital health interventions: Narrative review. Journal of Medical Internet Research, 26, e55130. https://doi.org/10.2196/55130

[30] Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2023). Capabilities of GPT-4 on medical challenge problems. arXiv preprint arXiv:2303.13375.

[31] Palaniappan, K., Lin, E. Y. T., & Vogel, S. (2024). Global Regulatory Frameworks for the Use of Artificial Intelligence (AI) in the Healthcare Services Sector. Healthcare (Basel, Switzerland), 12(5), 562.

[32] Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318). https://doi.org/10.3115/1073083.1073135

[33] Patwardhan, B., Mutalik, G., & Tillu, G. (2020). Integrative approaches for health: Biomedical research, Ayurveda and Yoga. Academic Press.

[34] Post, M. (2018). A call for clarity in reporting BLEU scores. In Proceedings of the Third Conference on Machine Translation: Research Papers (pp. 186-191).

[35] Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347-1358. https://doi.org/10.1056/NEJMra1814259

[36] Rao G. H. (2015). Integrative approach to health: Challenges and opportunities. Journal of Ayurveda and integrative medicine, 6(3), 215–219.

[37] Rasi, Sasan. (2020). Impact of Language Barriers on Access to Healthcare Services by Immigrant Patients: A systematic review. Asia-Pacific Journal of Health Management. 15. 35-48. 10.24083/apjhm.v15i1.271.

[38] Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp. 3982-3992). https://doi.org/10.18653/v1/D19-1410

[39] Schwalbe, N., & Wahl, B. (2020). Artificial intelligence and the future of global health. The Lancet, 395(10236), 1579-1586. https://doi.org/10.1016/S0140-6736(20)30226-9

[40] Sharma, R., Prakash, A., Chauhan, R., & Dhibar, D. P. (2021). Overcrowding an encumbrance for an emergency health-care system: A perspective of Health-care providers from tertiary care center in Northern India. Journal of education and health promotion, 10, 5. https://doi.org/10.4103/jehp.jehp_289_20

[41] Shortliffe, E. H. (1976). Computer-based medical consultations: MYCIN. Elsevier.

[42] Singh, P. K., Rai, R. K., Alagarajan, M., & Singh, L. (2022). Determinants of maternity care services utilization among married adolescents in rural India. PLoS One, 17(3), e0245468. https://doi.org/10.1371/journal.pone.0245468

[43] Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., ... Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172-180. https://doi.org/10.1038/s41586-023-06291-2

[44] Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, K., Pfohl, S., Cole-Lewis, H., Neal, D., Schaekermann, M., Wang, A., Amin, M., Lachgar, S., Mansfield, P., Prakash, S., Green, B., Dominowska, E., Arcas, B. A., ... Natarajan, V. (2022). Large language models encode clinical knowledge. arXiv preprint arXiv:2212.13138. https://doi.org/10.48550/arXiv.2212.13138

[45] Sterling, N. W., Patzer, R. E., Di, M., & Schrager, J. D. (2019). Prediction of emergency department patient disposition based on natural language processing of triage notes. International Journal of Medical Informatics, 129, 184-188.

[46] Victor, A. (2025, February 5). Artificial intelligence in global health: An unfair future for health in Sub-Saharan Africa? Health Affairs Scholar, 3(2), qxaf023. https://doi.org/10.1093/haschl/qxaf023

[47] Vithlani, J., Hawksworth, C., Elvidge, J., Ayiku, L., & Dawoud, D. (2023). Economic evaluations of artificial intelligence-based healthcare interventions: a systematic literature review of best practices in their conduct and reporting. Frontiers in pharmacology, 14, 1220950. https://doi.org/10.3389/fphar.2023.1220950

[48] Wahl, B., Cossy-Gantner, A., Germann, S., & Schwalbe, N. R. (2018). Artificial intelligence (AI) and global health: How can AI contribute to health in resource-poor settings? BMJ Global Health, 3(4), e000798. https://doi.org/10.1136/bmjgh-2018-000798

[49] WangAyers, J. W., Poliak, A., Dredze, M., Leas, E. C., Zhu, Z., Kelley, J. B., Faix, D. J., Goodman, A. M., Longhurst, C. A., Hogarth, M., & Smith, D. M. (2023). Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183(6), 589-596.

[50] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., & others. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824-24837.

[51] White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., & Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with ChatGPT. arXiv preprint arXiv:2302.11382. https://doi.org/10.48550/arXiv.2302.11382

[52] Wolff, J., Pauling, J., Keck, A., & Baumbach, J. (2020). The Economic Impact of Artificial Intelligence in Health Care: Systematic Review. Journal of medical Internet research, 22(2), e16866.

[53] Xie Q, Schenck EJ, Yang HS, Chen Y, Peng Y, Wang F. Faithful AI in Medicine: A Systematic Review with Large Language Models and Beyond. Res Sq [Preprint]. 2023 Dec 4:rs.3.rs-3661764. doi: 10.21203/rs.3.rs-3661764/v1. PMID: 38106170; PMCID: PMC10723541.

[54] Yang, X., Chen, A., PourNejatian, N., Shin, H. C., Smith, K. E., Parisien, C., Compas, C., Martin, C., Costa, A. B., Flores, M. G., Zhang, Y., Magoc, T., Harle, C. A., Lipori, G., Mitchell, D. A., Hogan, W. R., Shenkman, E. A., Bian, J., & Wu, Y. (2023). GatorTron: A large clinical language model to unlock patient information from unstructured electronic health records. npj Digital Medicine, 6, 115. https://doi.org/10.1038/s41746-023-00862-8

[55] Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.

How to cite this paper

Dr. P. A. Manoj Kumar, Dileep Parasu, Sarvesh Shashikumar, Kalyan Guru "Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings" Iconic Research And Engineering Journals Volume 9 Issue 3 2025 Page 1086-1109 https://doi.org/10.64388/IREV9I3-1710732-2258
Dr. P. A. Manoj Kumar, Dileep Parasu, Sarvesh Shashikumar, Kalyan Guru "Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025, doi: https://doi.org/10.64388/IREV9I3-1710732-2258
Dr. P. A. Manoj Kumar, Dileep Parasu, Sarvesh Shashikumar, Kalyan Guru (2025). Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings. Iconic Research And Engineering Journals, 9(3). doi: https://doi.org/10.64388/IREV9I3-1710732-2258
Dr. P. A. Manoj Kumar, Dileep Parasu, Sarvesh Shashikumar, Kalyan Guru "Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025. Crossref, https://doi.org/10.64388/IREV9I3-1710732-2258
@article{1710732,
      author = {Dr. P. A. Manoj Kumar, Dileep Parasu, Sarvesh Shashikumar, Kalyan Guru},
      title = {Generative AI and Strategic Prompt Engineering in Emergency Care: A Multi-Center Randomized Controlled Trial with Natural Language Processing Validation in Indian Healthcare Settings},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {3},
      pages = {1086-1109},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1710732.pdf},
      abstract = {Background: Indian emergency medicine is now under severe shortage of physicians, face a surge of patients and broad language barriers which require new AI methods that respond in the local health environments.

Study methods: In the study, we implemented a multi-center randomized controlled trial in several leading academic medical centers in India between 01 January 2023 and 31 December 2023. We did the study on 1,000 adult emergency department patients and assigned them to AI-assisted care (ChatGPT-4 with culturally-adapted prompt engineering) vs. standard care. Length of stay, adverse events at 30 days, and diagnostics accuracy were used as primary endpoints. The quality of AI-produced clinical summaries was measured with ROUGE, BLEU, LSA metrics and compared to the way it was documented by the physician. There was blindness to treatment in all the outcome assessors.

Findings: Of 1,000 randomized individuals (500 AI-assisted, 500 standard care) non-inferiority of AI-assisted care was shown in diagnostic accuracy (AI-assisted care 94.8%; standard care 94.2%; difference 0.6%, 95% CI: -2.1 to 3.3), and AI-assisted care had a superior performance in length of stay (AI-assisted care median 3.1; standard care median 4.3 hours; difference -1.2 hours). The Natural language processing evaluation showed high agreement, ROUGE-L scores 0.862?0.11, ROUGE-2F scores 0.804?0.14, and 689/1,000 (68.9%) cases scored 0.85 or above. The use of AI-assisted care saved physicians 38 percent of documentation time (P<0.001), raised clinical guidelines compliance by 23 percent (P<0.001) and raised patient satisfaction ratings (8.6 vs. 7.8; P<0.001). The cost-effectiveness analysis displayed savings of 2,847 Indian rupees per patient.

Conclusions: It was accomplished with excellent clinical results and outstanding financial cost savings, relative to the clinical outcomes, cultural adaptations of prompt engineering and AI-aided emergency care. These results confirm the use of AI nationwide in Indian emergency medicine.},
      keywords = {Artificial Intelligence, Emergency Medicine, Prompt Engineering, Natural Language Processing, ROUGE Score, Clinical Decision Support, Indian Healthcare, Cultural Adaptation},
      month = {September},
      doi = {https://doi.org/10.64388/IREV9I3-1710732-2258}
  }