Home / Current Issue / Paper 1719863
Vision Language Models (VLMs): A Comprehensive Review
Subject area: Science,Engineering and Technology · Area of research: Vision Language Models
Abstract
Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.
Keywords
Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.
How to cite this paper
@article{1719863,
author = {Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti},
title = {Vision Language Models (VLMs): A Comprehensive Review},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {8},
number = {6},
pages = {1329-1337},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1719863.pdf},
abstract = {Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.},
keywords = {Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.},
month = {December},
}