Enhancing Sentence Representation with Visually-supervised Multimodal Pre-training
Zhe Li, Laurence T. Yang, Xin Nie, Bocheng Ren, Xianjun Deng
摘要
Large-scale pre-trained language models have garnered significant attention in recent years due to their effectiveness in extracting sentence representations. However, most pre-trained models currently use transformer-based encoder with a single modality and are primarily designed for specific tasks such as natural language inference and question-answering. Unfortunately, this approach neglects the complementary information provided by multimodal data, which can enhance the effectiveness of sentence representation. To address this issue, we propose a Visually-supervised Pre-trained Multimodal Model (ViP) for sentence representation. Our model leverages diverse label-free multimodal proxy tasks to embed visual information into language, facilitating effective modality alignment and complementarity exploration. Additionally, our model utilizes a novel approach to distinguish highly similar negative and positive samples. We conduct comprehensive downstream experiments on natural language understanding and sentiment classification, demonstrating that ViP outperforms both existing unimodal and multimodal pre-trained models. Our contributions include a novel approach to multimodal pre-training and a state-of-the-art model for sentence representation that incorporates visual information.1 Our code is available at https://github.com/gentlefress/ViP
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive LearningZiqi Zhou, Shengshan Hu, Minghui Li, Hangtao Zhang 等ACM MM 2023 · 被引用 62 次
- From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent GuidanceZhe Li, Yangyang Wei, Boan Zhu, Yibo Peng 等ICLR 2026 · 被引用 29 次
- MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive LearningZhe Li, Laurence T. Yang, Bocheng Ren, Xin Nie 等CVPR 2024
- End-to-End Language-Action Model for Humanoid Whole Body ControlYuxuan Wang, Haobin Jiang, Shiqing Yao, Ziluo Ding 等CVPR 2026
相关 Paper
- CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language AlignmentHongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu 等ICLR 2023 · 被引用 53 次
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded SupervisionHao Tan, Mohit BansalEMNLP 2020 · 被引用 72 次
- End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingChi Chen, Peng Li, Maosong Sun, Yang LiuEMNLP 2022 · 被引用 7 次
- Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment AnalysisYan Ling, Jianfei Yu, Rui XiaACL 2022 · 被引用 116 次
- Sentiment-Aware Word and Sentence Level Pre-training for Sentiment AnalysisShuai Fan, Chen Lin, Haonan Li, Zhenghao Lin 等EMNLP 2022 · 被引用 20 次
