Home / Current Issue / Paper 1713739
Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models
Subject area: Science,Engineering and Technology · Area of research: Multimodal Vision Language
DOI: https://doi.org/10.64388/IREV9I7-1713739
Abstract
Social media has officially crossed the 5-billion-user mark, and we are on track to hit 6 billion within just a few years. Over 90% of teens and the majority of children now bypass text and photos in favor of fast-paced, short-form video. This shift has put immense pressure on platforms to keep up. When you?re dealing with millions of uploads every single day, traditional moderation tools which often look at data in silos just can?t keep pace with the nuance and context of modern content. To solve this, the industry is moving toward more sophisticated AI models, like Multimodal Large Language Models (MLLMs). These systems are designed to "think" more like humans by processing video, sound, and text simultaneously to catch harmful content that older models might miss. Implementing a service for filtering YouTube Shorts and Instagram Reels requires a sophisticated architectural design that balances information integrity with computational speed.
How to cite this paper
@article{1713739,
author = {Niharika Patidar, Dr. Sachin Patel},
title = {Implementation Architectures and Systematic Survey of Multimodal Vision-Language Models},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {7},
pages = {1783-1790},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1713739.pdf},
abstract = {Social media has officially crossed the 5-billion-user mark, and we are on track to hit 6 billion within just a few years. Over 90% of teens and the majority of children now bypass text and photos in favor of fast-paced, short-form video. This shift has put immense pressure on platforms to keep up. When you?re dealing with millions of uploads every single day, traditional moderation tools which often look at data in silos just can?t keep pace with the nuance and context of modern content. To solve this, the industry is moving toward more sophisticated AI models, like Multimodal Large Language Models (MLLMs). These systems are designed to "think" more like humans by processing video, sound, and text simultaneously to catch harmful content that older models might miss. Implementing a service for filtering YouTube Shorts and Instagram Reels requires a sophisticated architectural design that balances information integrity with computational speed.},
month = {January},
doi = {https://doi.org/10.64388/IREV9I7-1713739}
}