PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models
Yuan Yao, Qianyu Chen, Ao Zhang, Wei Ji, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun
摘要
Vision-language pre-training (VLP) has shown impressive performance on a wide range of cross-modal tasks, where VLP models without reliance on object detectors are becoming the mainstream due to their superior computation efficiency and competitive performance. However, the removal of object detectors also deprives the capability of VLP models in explicit object modeling, which is essential to various position-sensitive vision-language (VL) tasks, such as referring expression comprehension and visual commonsense reasoning. To address the challenge, we introduce PEVL that enhances the pre-training and prompt tuning of VLP models with explicit object position modeling. Specifically, PEVL reformulates discretized object positions and language in a unified language modeling framework, which facilitates explicit VL alignment during pre-training, and also enables flexible prompt tuning for various downstream tasks. We show that PEVL enables state-of-the-art performance of detector-free VLP models on position-sensitive tasks such as referring expression comprehension and phrase grounding, and also improves the performance on position-insensitive tasks with grounded inputs. We make the data and code for this paper publicly available at https://github.com/thunlp/PEVL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi 等ICLR 2023 · 被引用 110 次
- NExT-Chat: An LMM for Chat, Detection and SegmentationAo Zhang, Yuan Yao, Wei Ji, Zhiyuan Liu 等ICML 2024 · 被引用 85 次
- Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCheng Cheng, Lin Song, Ruoyi Xue, Hang Wang 等NeurIPS 2023 · 被引用 65 次
- Panoptic Scene Graph Generation with Semantics-Prototype LearningLi Li, Wei Ji, Yiming Wu, Mengze Li 等AAAI 2024 · 被引用 63 次
- Voila-A: Aligning Vision-Language Models with User's Gaze AttentionKun Yan, Zeyu Wang, Lei Ji, Yuntao Wang 等NeurIPS 2024 · 被引用 43 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
相关 Paper
- Position-Guided Text Prompt for Vision-Language Pre-TrainingJinpeng Wang, Pan Zhou, Mike Zheng Shou, Shuicheng YanCVPR 2023
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi 等ACL 2021
- End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression MatchingChi Chen, Peng Li, Maosong Sun, Yang LiuEMNLP 2022 · 被引用 7 次
- Align and Prompt: Video-and-Language Pre-training with Entity PromptsDongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles 等CVPR 2022
- Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentMingyang Zhou, Licheng Yu, Amanpreet Singh, Mengjiao Wang 等CVPR 2022 · 被引用 29 次
