Linguistic-Aware Patch Slimming Framework for Fine-Grained Cross-Modal Alignment
Zheren Fu, Lei Zhang, Hou Xia, Zhendong Mao
摘要
Cross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities between images and texts. Traditional fine-grained alignment methods heavily rely on pre-trained object detectors to extract region features for subsequent region-word alignment, thereby incurring substantial computational costs for region detection and error propagation issues for two-stage training. In this paper, we focus on the mainstream vision transformer, incorporating patch features for patch-word alignment, while addressing the resultant issue of visual patch redundancy and patch ambiguity for semantic alignment. We propose a novel Linguistic-Aware Patch Slimming (LAPS) framework for fine-grained alignment, which explicitly identifies redundant visual patches with language supervision and rectifies their semantic and spatial information to facilitate more effective and consistent patchword alignment. Extensive experiments on various evaluation benchmarks and model backbones show LAPS outperforms the state-of-the-art fine-grained alignment methods by 5%-15% rSum. Our code is available at https: //github.com/CrossmodalGroup/LAPS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal RetrievalLanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu 等NeurIPS 2025 · 被引用 12 次
- Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language AlignmentYang Liu, Mengyuan Liu, Shudong Huang, Jiancheng LvAAAI 2025 · 被引用 8 次
- Homology Consistency Constrained Efficient Tuning for Vision-Language ModelsHuatian Zhang, Lei Zhang, Yongdong Zhang, Zhendong MaoNeurIPS 2024 · 被引用 5 次
- Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text MatchingYang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang 等ICCV 2025 · 被引用 3 次
- Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language ModelsYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal AlignmentXinyu Mao, Junsi Li, Haoji Zhang, Yu Liang 等ICML 2026
- COPA : Efficient Vision-Language Pre-training through Collaborative Object- and Patch-Text AlignmentChaoya Jiang, Haiyang Xu, Wei Ye, Qinghao Ye 等ACM MM 2023 · 被引用 10 次
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng 等CVPR 2026
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language ModelsJiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning 等NeurIPS 2025 · 被引用 8 次
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang 等NeurIPS 2025 · 被引用 7 次
