Home / Current Issue / Paper 1722541
Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation
Subject area: Science,Engineering and Technology · Area of research: Multimodal Deep Learning
Abstract
We propose a unified multimodal transformer with a mixture-of-experts (MoE) core for joint semantic segmentation and caption generation. A ConvNeXt-B visual encoder and a T5 text encoder feed a shared MoE block whose experts specialize by modality and task, routed dynamically per token. On LVIS and NYU-Depth v2 the model attains 70.9% mIoU and a BLEU-4 of 34, improving over a strong generative baseline while keeping active compute modest through sparse routing.
How to cite this paper
Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals Volume 8 Issue 3 2024 Page 1179-1184
Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals, vol. 8, no. 3, Sep. 2024
Olivia Brennan, Karthik Venkat, Lucas Moreau (2024). Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation. Iconic Research And Engineering Journals, 8(3).
Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals, vol. 8, no. 3, Sep. 2024.
@article{1722541,
author = {Olivia Brennan, Karthik Venkat, Lucas Moreau},
title = {Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {8},
number = {3},
pages = {1179-1184},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722541.pdf},
abstract = {We propose a unified multimodal transformer with a mixture-of-experts (MoE) core for joint semantic segmentation and caption generation. A ConvNeXt-B visual encoder and a T5 text encoder feed a shared MoE block whose experts specialize by modality and task, routed dynamically per token. On LVIS and NYU-Depth v2 the model attains 70.9% mIoU and a BLEU-4 of 34, improving over a strong generative baseline while keeping active compute modest through sparse routing.},
month = {September},
}