International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1719863

1719863PublishedVol 8 · Issue 6

Vision Language Models (VLMs): A Comprehensive Review

Anmol Diya Goell Dhruv Jain Kushal Preeti

Subject area: Science,Engineering and Technology  ·  Area of research: Vision Language Models

Abstract

Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.

Keywords

Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.

How to cite this paper

Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti "Vision Language Models (VLMs): A Comprehensive Review" Iconic Research And Engineering Journals Volume 8 Issue 6 2024 Page 1329-1337
Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti "Vision Language Models (VLMs): A Comprehensive Review" Iconic Research And Engineering Journals, vol. 8, no. 6, Dec. 2024
Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti (2024). Vision Language Models (VLMs): A Comprehensive Review. Iconic Research And Engineering Journals, 8(6).
Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti "Vision Language Models (VLMs): A Comprehensive Review" Iconic Research And Engineering Journals, vol. 8, no. 6, Dec. 2024.
@article{1719863,
      author = {Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti},
      title = {Vision Language Models (VLMs): A Comprehensive Review},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {8},
      number = {6},
      pages = {1329-1337},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1719863.pdf},
      abstract = {Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.},
      keywords = {Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.},
      month = {December},
  }