ROSITA: Enhancing Vision-and-Language Semantic Alignments via Cross- and Intra-modal Knowledge Integration
Yuhao Cui, Zhou Yu, Chunqi Wang, Zhongzhou Zhao, Ji Zhang, Meng Wang, Jun Yu
Abstract
Vision-and-language pretraining (VLP) aims to learn generic multimodal representations from massive image-text pairs. While various successful attempts have been proposed, learning fine-grained semantic alignments between image-text pairs plays a key role in their approaches. Nevertheless, most existing VLP approaches have not fully utilized the intrinsic knowledge within the image-text pairs, which limits the effectiveness of the learned alignments and further restricts the performance of their models. To this end, we introduce a new VLP method called ROSITA, which integrates the cross- and intra-modal knowledge in a unified scene graph to enhance the semantic alignments. Specifically, we introduce a novel structural knowledge masking (SKM) strategy to use the scene graph structure as a priori to perform masked language (region) modeling, which enhances the semantic alignments by eliminating the interference information within and across modalities. Extensive ablation studies and comprehensive analysis verifies the effectiveness of ROSITA in semantic alignments. Pretrained with both in-domain and out-of-domain datasets, ROSITA significantly outperforms existing state-of-the-art VLP methods on three typical vision-and-language tasks over six benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7784d0d1-5efe-491b-8407-d3772c630f3dCited by top-tier papers6
- Align, Reason and Learn: Enhancing Medical Vision-and-Language Pre-training with KnowledgeZhihong Chen, Guanbin Li, Xiang WanACM MM 2022 · 82 citations
- Multi-Attention Network for Compressed Video Referring Object SegmentationWeidong Chen, Dexiang Hong, Yuankai Qi, Zhenjun Han et al.ACM MM 2022 · 49 citations
- MVPTR: Multi-Level Semantic Alignment for Vision-Language Pre-Training via Multi-Stage LearningZejun Li, Zhihao Fan, Huaixiao Tou, Jingjing Chen et al.ACM MM 2022 · 15 citations
- Counterfactually Measuring and Eliminating Social Bias in Vision-Language Pre-training ModelsYi Zhang, Junyang Wang, Jitao SangACM MM 2022 · 11 citations
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 8 citations
Builds on9
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang et al.AAAI 2020 · 898 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 391 citations
Related papers
- Seeing What You Miss: Vision-Language Pre-training with Semantic Completion LearningYatai Ji, Rongcheng Tu, Jie Jiang, Weijie Kong et al.CVPR 2023
- ViLTA: Enhancing Vision-Language Pre-training through Textual AugmentationWeihan Wang, Zhen Yang, Bin Xu, Juanzi Li et al.ICCV 2023 · 11 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang et al.CVPR 2022 · 38 citations
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi et al.ACL 2021
