Home / Current Issue / Paper 1722526
Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer
Subject area: Science,Engineering and Technology · Area of research: Bidirectional Vision-Language Generation
Abstract
We present a unified cross-modal transformer that generates in both directions between images and language: producing semantic segmentation masks from captions and captions from masks within a single model. A shared cross-modal attention stack, built on a Swin-B vision encoder and a BERT text encoder, aligns the two modalities in a common space. On Cityscapes and PASCAL-Context the model reaches 69.7% mIoU for text-to-mask and a BLEU-4 of 33 for mask-to-text, outperforming strong unidirectional baselines.
How to cite this paper
Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals Volume 7 Issue 5 2023 Page 475-480
Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals, vol. 7, no. 5, Nov. 2023
Camille Dubois, Arjun Nair, Sofia Petrova (2023). Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer. Iconic Research And Engineering Journals, 7(5).
Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals, vol. 7, no. 5, Nov. 2023.
@article{1722526,
author = {Camille Dubois, Arjun Nair, Sofia Petrova},
title = {Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer},
journal = {Iconic Research And Engineering Journals},
year = {2023},
volume = {7},
number = {5},
pages = {475-480},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722526.pdf},
abstract = {We present a unified cross-modal transformer that generates in both directions between images and language: producing semantic segmentation masks from captions and captions from masks within a single model. A shared cross-modal attention stack, built on a Swin-B vision encoder and a BERT text encoder, aligns the two modalities in a common space. On Cityscapes and PASCAL-Context the model reaches 69.7% mIoU for text-to-mask and a BLEU-4 of 33 for mask-to-text, outperforming strong unidirectional baselines.},
month = {November},
}