International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1713739

1713739PublishedVol 9 · Issue 7

Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models

Niharika Patidar Dr. Sachin Patel

Subject area: Science,Engineering and Technology  ·  Area of research: Multimodal Vision Language

DOI: https://doi.org/10.64388/IREV9I7-1713739

Abstract

Social media has officially crossed the 5-billion-user mark, and we are on track to hit 6 billion within just a few years. Over 90% of teens and the majority of children now bypass text and photos in favor of fast-paced, short-form video. This shift has put immense pressure on platforms to keep up. When you?re dealing with millions of uploads every single day, traditional moderation tools which often look at data in silos just can?t keep pace with the nuance and context of modern content. To solve this, the industry is moving toward more sophisticated AI models, like Multimodal Large Language Models (MLLMs). These systems are designed to "think" more like humans by processing video, sound, and text simultaneously to catch harmful content that older models might miss. Implementing a service for filtering YouTube Shorts and Instagram Reels requires a sophisticated architectural design that balances information integrity with computational speed.

How to cite this paper

Niharika Patidar, Dr. Sachin Patel "Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models" Iconic Research And Engineering Journals Volume 9 Issue 7 2026 Page 1783-1790 https://doi.org/10.64388/IREV9I7-1713739
Niharika Patidar, Dr. Sachin Patel "Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models" Iconic Research And Engineering Journals, vol. 9, no. 7, Jan. 2026, doi: https://doi.org/10.64388/IREV9I7-1713739
Niharika Patidar, Dr. Sachin Patel (2026). Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models. Iconic Research And Engineering Journals, 9(7). doi: https://doi.org/10.64388/IREV9I7-1713739
Niharika Patidar, Dr. Sachin Patel "Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models" Iconic Research And Engineering Journals, vol. 9, no. 7, Jan. 2026. Crossref, https://doi.org/10.64388/IREV9I7-1713739
@article{1713739,
      author = {Niharika Patidar, Dr. Sachin Patel},
      title = {Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {7},
      pages = {1783-1790},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1713739.pdf},
      abstract = {Social media has officially crossed the 5-billion-user mark, and we are on track to hit 6 billion within just a few years. Over 90% of teens and the majority of children now bypass text and photos in favor of fast-paced, short-form video. This shift has put immense pressure on platforms to keep up. When you?re dealing with millions of uploads every single day, traditional moderation tools which often look at data in silos just can?t keep pace with the nuance and context of modern content. To solve this, the industry is moving toward more sophisticated AI models, like Multimodal Large Language Models (MLLMs). These systems are designed to "think" more like humans by processing video, sound, and text simultaneously to catch harmful content that older models might miss. Implementing a service for filtering YouTube Shorts and Instagram Reels requires a sophisticated architectural design that balances information integrity with computational speed.},
      month = {January},
      doi = {https://doi.org/10.64388/IREV9I7-1713739}
  }