International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722526

1722526 Vol 7 · Issue 5 Download Paper

Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer

Camille Dubois Arjun Nair Sofia Petrova

Subject area: Science,Engineering and Technology  ·  Area of research: Bidirectional Vision-Language Generation

Abstract

We present a unified cross-modal transformer that generates in both directions between images and language: producing semantic segmentation masks from captions and captions from masks within a single model. A shared cross-modal attention stack, built on a Swin-B vision encoder and a BERT text encoder, aligns the two modalities in a common space. On Cityscapes and PASCAL-Context the model reaches 69.7% mIoU for text-to-mask and a BLEU-4 of 33 for mask-to-text, outperforming strong unidirectional baselines.

How to cite this paper

Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals Volume 7 Issue 5 2023 Page 475-480
Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals, vol. 7, no. 5, Nov. 2023
Camille Dubois, Arjun Nair, Sofia Petrova (2023). Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer. Iconic Research And Engineering Journals, 7(5).
Camille Dubois, Arjun Nair, Sofia Petrova "Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer" Iconic Research And Engineering Journals, vol. 7, no. 5, Nov. 2023.
@article{1722526,
      author = {Camille Dubois, Arjun Nair, Sofia Petrova},
      title = {Bidirectional Vision-Language Generation: Synthesizing Segmentation Masks and Captions with a Unified Cross-Modal Transformer},
      journal = {Iconic Research And Engineering Journals},
      year = {2023},
      volume = {7},
      number = {5},
      pages = {475-480},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722526.pdf},
      abstract = {We present a unified cross-modal transformer that generates in both directions between images and language: producing semantic segmentation masks from captions and captions from masks within a single model. A shared cross-modal attention stack, built on a Swin-B vision encoder and a BERT text encoder, aligns the two modalities in a common space. On Cityscapes and PASCAL-Context the model reaches 69.7% mIoU for text-to-mask and a BLEU-4 of 33 for mask-to-text, outperforming strong unidirectional baselines.},
      month = {November},
  }