Home / Current Issue / Paper 1718331
Personalized Image-to-Audio Bedtime Story Generation Using Multimodal AI with User Profiling
Subject area: Science,Engineering and Technology · Area of research: Multimodal AI and Personalized Storytelling
DOI: https://doi.org/10.64388/IREV9I11-1718331
Abstract
Bedtime stories help children improve imagination, communication skills, and emotional connection with parents. However, generating personalized and engaging bedtime stories daily can be challenging for parents. Existing AI-based storytelling systems mainly rely on text inputs, resulting in generic narratives that lack contextual relevance and emotional adaptation. This paper presents a multimodal framework that transforms an input image into a personalized bedtime audio story using user profile attributes such as age, preferences, and mood. The system combines image understanding, generative language models, and text-to-speech techniques to produce context-aware and emotionally adaptive narratives with calming themes. The proposed approach improves storytelling by integrating visual context and personalization, making stories more engaging and meaningful. This work highlights the potential of multimodal AI in enhancing bedtime routines and supporting child well-being.
Keywords
Multimodal AI, Personalized Storytelling, Image-to-Audio, Text-to-Speech, User Profiling
References
[1] S. Sharma, A. Verma, and R. Singh, “AI-Based Story Generation Using Natural Language Processing,” International Journal of Computer Applications, vol. 183, no. 42, pp. 15–20, 2021.
[2] R. Verma and S. Bansal, “Sequence Modeling Techniques for Story Generation,” in Proceedings of the IEEE International Conference on Artificial Intelligence, 2023, pp. 120–125.
[3] A. Kulkarni and S. Chatterjee, “Image-to-Text Generation Using Deep Learning,” International Journal of Computer Vision, vol. 130, no. 5, pp. 1120–1135, 2022.
[4] K. Rao and A. Mishra, “Neural Text-to-Speech Systems: A Review,” IEEE Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 230–245, 2021.
[5] A. Reddy and R. Banerjee, “User Profiling for Personalized Recommendation Systems,” Journal of Intelligent Systems, vol. 30, no. 2, pp. 210–220, 2021.
[6] T. Brown et al., “Language Models are Few-Shot Learners,” in Advances in Neural Information
[7] A. Vaswani et al., “Attention Is All You Need,” in Advances in Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008.
[8] Y. LeCun, Y. Bengio, and G. Hinton, “Deep Learning,” Nature, vol. 521, no. 7553, pp. 436–444, 2015.
[9] D. Jurafsky and J. H. Martin, Speech and Language Processing, 3rd ed., Pearson, 2022.
[10] A. Ahmed and P. Kapoor, “Multimodal AI for Content Generation: A Comprehensive Study,” IEEE Access, vol. 12, pp. 34567–34580, 2024.
How to cite this paper
@article{1718331,
author = {Krutika Sushil Nikumbh, Dr. Prakash Kene},
title = {Personalized Image-to-Audio Bedtime Story Generation Using Multimodal AI with User Profiling},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {11},
pages = {4443-4449},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718331.pdf},
abstract = {Bedtime stories help children improve imagination, communication skills, and emotional connection with parents. However, generating personalized and engaging bedtime stories daily can be challenging for parents. Existing AI-based storytelling systems mainly rely on text inputs, resulting in generic narratives that lack contextual relevance and emotional adaptation. This paper presents a multimodal framework that transforms an input image into a personalized bedtime audio story using user profile attributes such as age, preferences, and mood. The system combines image understanding, generative language models, and text-to-speech techniques to produce context-aware and emotionally adaptive narratives with calming themes. The proposed approach improves storytelling by integrating visual context and personalization, making stories more engaging and meaningful. This work highlights the potential of multimodal AI in enhancing bedtime routines and supporting child well-being.},
keywords = {Multimodal AI, Personalized Storytelling, Image-to-Audio, Text-to-Speech, User Profiling},
month = {May},
doi = {https://doi.org/10.64388/IREV9I11-1718331}
}