Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
Yeyuan Wang, Dehong Gao, Lei Yi, Linbo Jin, Jinxia Zhang, Libin Yang, Xiaoyan Cai
摘要
Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many nuanced vision-language applications, remains limited. Prevailing VLP models often overlook the intricate distinctions in expressing different modal features and typically depend on the similarity of holistic features for cross-modal interactions. Moreover, these models directly align and integrate features from different modalities, focusing more on coarse-grained general representations, thus failing to capture the nuanced differences necessary for tasks demanding a more detailed perception. In response to these limitations, we introduce Negative Augmented Samples(NAS), a refined vision-language pretraining model that innovatively incorporates NAS to specifically address the challenge of fine-grained understanding. NAS utilizes a Visual Dictionary(VD) as a semantic bridge between visual and linguistic domains. Additionally, it employs a Negative Visual Augmentation(NVA) method based on the VD to generate challenging negative image samples. These samples deviate from positive samples exclusively at the token level, thereby necessitating that the model discerns the subtle disparities between positive and negative samples with greater precision. Comprehensive experiments validate the efficacy of NAS components and underscore its potential to enhance fine-grained vision-language comprehension.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long CaptionsZiteng Wang, Siqi Yang, Limeng Qiao, Lin MaNeurIPS 2025 · 被引用 5 次
- Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive InteractionYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026
- Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language ModelsJia-Wei Hai, Yijun Wang, Xiu-Shen WeiICML 2026
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentAhtamjan Ahmat, Lei Wang, Yating Yang, Bo Ma 等WWW 2025 · 被引用 2 次
- ViLTA: Enhancing Vision-Language Pre-training through Textual AugmentationWeihan Wang, Zhen Yang, Bin Xu, Juanzi Li 等ICCV 2023 · 被引用 11 次
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang 等ICCV 2023 · 被引用 6 次
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsYan Zeng, Xinsong Zhang, Hang LiICML 2022 · 被引用 371 次
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 被引用 16 次
