Home / Current Issue / Paper 1719863
Vision Language Models (VLMs): A Comprehensive Review
Subject area: Science,Engineering and Technology · Area of research: Vision Language Models
Abstract
Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.
Keywords
Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.
References
[1] J. Achiam et al., "GPT-4 Technical Report," arXiv preprint arXiv:2303.08774, 2023.
[2] OpenAI, "GPT-4V(ision) System Card," OpenAI Technical Report, 2023.
[3] H. Liu, C. Li, Q. Wu, and Y. J. Lee, "Visual Instruction Tuning," in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
[4] H. Liu, C. Li, Y. Li, and Y. J. Lee, "Improved Baselines with Visual Instruction Tuning," arXiv preprint arXiv:2310.03744, 2023.
[5] J. Bai et al., "Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities," arXiv preprint arXiv:2308.12966, 2023.
[6] L. Yuan et al., "Florence: A New Foundation Model for Computer Vision," arXiv preprint arXiv:2111.11432, 2021.
[7] A. Radford et al., "Learning Transferable Visual Models From Natural Language Supervision," in Proc. Int. Conf. on Machine Learning (ICML), 2021.
[8] J. Li, D. Li, S. Savarese, and S. Hoi, "BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models," in Proc. ICML, 2023.
[9] A. Dosovitskiy et al., "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale," in Proc. Int. Conf. on Learning Representations (ICLR), 2021.
[10] Z. Liu et al., "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows," in Proc. IEEE/CVF Int. Conf. on Computer Vision (ICCV), 2021.
[11] Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh, "Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering," in Proc. CVPR, 2017.
[12] A. Singh et al., "Towards VQA Models That Can Read," in Proc. CVPR, 2019.
[13] Y. Liu et al., "MMBench: Is Your Multi-modal Model an All-around Player?," arXiv preprint arXiv:2307.06281, 2023.
[14] T.-Y. Lin et al., "Microsoft COCO: Common Objects in Context," in Proc. European Conf. on Computer Vision (ECCV), 2014.
[15] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "ImageNet: A Large-Scale Hierarchical Image Database," in Proc. CVPR, 2009.
[16] J.-B. Alayrac et al., "Flamingo: A Visual Language Model for Few-Shot Learning," in Proc. NeurIPS, 2022.
[17] M. Mathew, D. Karatzas, and C. V. Jawahar, "DocVQA: A Dataset for VQA on Document Images," in Proc. IEEE Winter Conf. on Applications of Computer Vision (WACV), 2021.
[18] K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi, "OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge," in Proc. CVPR, 2019.
How to cite this paper
@article{1719863,
author = {Anmol, Diya Goell, Dhruv Jain, Kushal, Preeti},
title = {Vision Language Models (VLMs): A Comprehensive Review},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {8},
number = {6},
pages = {1329-1337},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1719863.pdf},
abstract = {Vision-Language Models (VLMs) represent one of the most significant advances in multimodal artificial intelligence, enabling systems to jointly reason over visual and textual information. This review surveys four representative VLMs — GPT-4 Vision, LLaVA, Qwen-VL, and Florence — examining their architectures, training paradigms, and downstream capabilities. We compare the models along dimensions of vision encoder design, language backbone, training data scale, and benchmark performance, and we discuss common failure modes such as hallucination, weak spatial grounding, and limited fine-grained visual reasoning. We further outline open challenges including data efficiency, evaluation standardization, and computational cost, and we propose promising directions such as unified tokenization, retrieval-augmented multimodal reasoning, and efficient adapter-based fine-tuning. This review is intended as a structured reference for researchers and practitioners entering the field of vision-language modeling.},
keywords = {Vision-Language Models, GPT-4V, LLaVA, Qwen-VL, Florence, multimodal learning, visual instruction tuning, transformers.},
month = {December},
}