International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722541

1722541 Vol 8 · Issue 3 Download Paper

Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation

Olivia Brennan Karthik Venkat Lucas Moreau

Subject area: Science,Engineering and Technology  ·  Area of research: Multimodal Deep Learning

Abstract

We propose a unified multimodal transformer with a mixture-of-experts (MoE) core for joint semantic segmentation and caption generation. A ConvNeXt-B visual encoder and a T5 text encoder feed a shared MoE block whose experts specialize by modality and task, routed dynamically per token. On LVIS and NYU-Depth v2 the model attains 70.9% mIoU and a BLEU-4 of 34, improving over a strong generative baseline while keeping active compute modest through sparse routing.

How to cite this paper

Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals Volume 8 Issue 3 2024 Page 1179-1184
Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals, vol. 8, no. 3, Sep. 2024
Olivia Brennan, Karthik Venkat, Lucas Moreau (2024). Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation. Iconic Research And Engineering Journals, 8(3).
Olivia Brennan, Karthik Venkat, Lucas Moreau "Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation" Iconic Research And Engineering Journals, vol. 8, no. 3, Sep. 2024.
@article{1722541,
      author = {Olivia Brennan, Karthik Venkat, Lucas Moreau},
      title = {Unified Multimodal Transformers for Joint Semantic Segmentation and Caption Generation},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {8},
      number = {3},
      pages = {1179-1184},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722541.pdf},
      abstract = {We propose a unified multimodal transformer with a mixture-of-experts (MoE) core for joint semantic segmentation and caption generation. A ConvNeXt-B visual encoder and a T5 text encoder feed a shared MoE block whose experts specialize by modality and task, routed dynamically per token. On LVIS and NYU-Depth v2 the model attains 70.9% mIoU and a BLEU-4 of 34, improving over a strong generative baseline while keeping active compute modest through sparse routing.},
      month = {September},
  }