International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1704609

1704609 Vol 6 · Issue 12 Download Paper

Generating Captions for Images Using Neural Networks

Ishaan Taneja Sunil Maggu

Subject area: Science,Engineering and Technology  ·  Area of research: Convolutional Neural Networks

Abstract

Different concepts in the field of Artificial Intelligence are on the rise these days, generating captions from given images, being one of them. The ability to train a machine to be provided with an image and then it being able to describe the details around the same can be used in various applications, be it robotics, or other businesses. The primary purpose of this paper is to recommend a model which describes images and provides its captions using concepts of Deep Learning and Machine translation. The model aims to detect different types of objects around an image, recognize the relationships between them and then generate the desired captions. The model, developed in Python, is trained using the Flickr 8K dataset in order to accomplish the same. The model was developed using Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM). In addition to discussing VGG16, a variant of CNN that has proven useful in our use case, the paper delves deeply into the fundamental notions of CNN. The wider use of this model is to help the masses, wherein, it can be used in image indexing to help those with visual impairments, also it can be implemented on some social networks and can be used in other applications as well.

References

[1] Haoran Wang ,Yue Zhang and Xiaosheng Yu, An Overview of Image Caption Generation Methods, ,2020

[2] Sequence to sequence -video to text by Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko,2017.

[3] ” Every Picture Tells a Story: Generating Sentences from Images.” Computer Vision ECCV (2016) by Farhadi, Ali, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hocken-maier, and David Forsyth

[4] Show and Tell: A Neural Image Caption Generator by Oriol Vinyal, Alexander Toshev, Samy Bengio, Dumitru Erhan, IEEE (2015)

[5] Krishnakumar, K.Kousalya, S.Gokul, R.Karthikeyan D.Kaviyarasu , Image Caption Generator using Deep Learning, International Journal of Advanced Science and Technology, Vol. 29, No. 3s, (2020), pp. 975-980ISSN: 2005-4238 IJAST

[6] P. Aishwarya Naidu1, Satvik Vats,Gehna Anand, Nalina V, A Deep Learning Model for Image Caption Generation, Published: 30/June/2020

[7] Rennie, Steven & Marcheret, E. & Mroueh, Youssef & Ross, Jarret & Goel, Vaibhava. (2017). Self-Critical Sequence Training for Image Captioning.

[8] Alzubaidi, L., Zhang, J., Humaidi, A.J. et al. Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. J Big Data 8, 53 (2021). https://doi.org/10.1186/s40537-021-00444-8

[9] Pranay Mathur, Aman Gill, Aayush Yadav, Anurag Mishra, and Nand Kumar Bansode,“Camera2Caption: A Real-Time Image Caption Generator”, International Conference on Computational Intelligence in Data Science(ICCIDS) - 2017

[10] Oriol Vinyals, Alexander Toshev, Samy Bengio, Dumitru Erhan, Show and Tell: A Neural Image Caption Generator. CVPR2015

[11] D. Bahdanau, K. Cho, and Y. Bengio. “Neural machine translation by jointly learning to align and translate. arXiv:1409.0473”, 2014.

[12] BLEU: A method for automatic evaluation of machine translation. InACL, 2002 by K. Papineni, S. Roukos, T. Ward, and W. J. Zhu.

[13] Priyanka Kalena, Nishi Malde, Aromal Nair, Saurabh Parkar, and Grishma Sharma, “Visual Image Caption Generator Using Deep Learning”, (ICAST-2019)

How to cite this paper

Ishaan Taneja, Sunil Maggu "Generating Captions for Images Using Neural Networks" Iconic Research And Engineering Journals Volume 6 Issue 12 2023 Page 214-218
Ishaan Taneja, Sunil Maggu "Generating Captions for Images Using Neural Networks" Iconic Research And Engineering Journals, vol. 6, no. 12, Jun. 2023
Ishaan Taneja, Sunil Maggu (2023). Generating Captions for Images Using Neural Networks. Iconic Research And Engineering Journals, 6(12).
Ishaan Taneja, Sunil Maggu "Generating Captions for Images Using Neural Networks" Iconic Research And Engineering Journals, vol. 6, no. 12, Jun. 2023.
@article{1704609,
      author = {Ishaan Taneja, Sunil Maggu},
      title = {Generating Captions for Images Using Neural Networks},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {6},
      number = {12},
      pages = {214-218},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1704609.pdf},
      abstract = {Different concepts in the field of Artificial Intelligence are on the rise these days, generating captions from given images, being one of them. The ability to train a machine to be provided with an image and then it being able to describe the details around the same can be used in various applications, be it robotics, or other businesses. The primary purpose of this paper is to recommend a model which describes images and provides its captions using concepts of Deep Learning and Machine translation. The model aims to detect different types of objects around an image, recognize the relationships between them and then generate the desired captions. The model, developed in Python, is trained using the Flickr 8K dataset in order to accomplish the same. The model was developed using Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM). In addition to discussing VGG16, a variant of CNN that has proven useful in our use case, the paper delves deeply into the fundamental notions of CNN. The wider use of this model is to help the masses, wherein, it can be used in image indexing to help those with visual impairments, also it can be implemented on some social networks and can be used in other applications as well.},
      month = {June},
  }