Rethinking Causal Mask Attention for Vision-Language Inference
Xiaohuan Pei, Tao Huang, Yanxiang Ma, Chang Xu
摘要
Causal attention has become a foundational mechanism in autoregressive Vision-Language models (VLMs), unifying textual and visual inputs under a single generative framework. However, existing causal mask-based strategies are inherited from large language models (LLMs) where they are tailored for text-only decoding, and their adaptation to vision tokens is insufficiently addressed in the prefill stage. Strictly masking future positions for vision queries introduces overly rigid constraints, which hinder the model’s ability to leverage future context that often contains essential semantic cues for accurate inference. In this work, we empirically investigate how different causal masking strategies affect vision-language inference and then propose a family of future-aware attentions tailored for this setting. We first empirically analyze the effect of previewing future tokens for vision queries and demonstrate that rigid masking undermines the model’s capacity to capture useful contextual semantic representations. Based on these findings, we propose a lightweight attention family that aggregates future visual context into past representations via pooling, effectively preserving the autoregressive structure while enhancing cross-token dependencies. We evaluate a range of causal masks across diverse vision-language inference settings and show that selectively compressing future semantic context into past representations benefits the inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Aligning What Vision-Language Models See and Perceive with Adaptive Information FlowChengxin Liu, Wonseok Choi, Chenshuang Zhang, Tae-Hyun OhCVPR 2026 · 被引用 2 次
- HInT: Hypergraph Infusion at the Structural Layers Improves Table UnderstandingWonjin Lee, Soomi Jeong, Kwang In KimICML 2026
它引用的顶会 Paper18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
相关 Paper
- -Attn: Decomposed Attention for Large Vision-and-Language ModelsChia-Wen Kuo, Sijie Zhu, Fan Chen, Xiaohui Shen 等ICCV 2025 · 被引用 1 次
- LearnPruner: Rethinking Attention-based Token Pruning in Vision Language ModelsRinyoichi Takezoe, Yaqian Li, Zi-Hao Bo, Anzhou Hou 等ICLR 2026 · 被引用 8 次
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal DecodingFeilong Tang, Chengzhi Liu, Zhongxing Xu, Ming Hu 等CVPR 2025
- A-VL: Adaptive Attention for Large Vision-Language ModelsJunyang Zhang, Mu Yuan, Ruiguang Zhong, Puhan Luo 等AAAI 2025 · 被引用 6 次
- Causal Graphical Models for Vision-Language Compositional UnderstandingFiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto, Lorenzo Baraldi 等ICLR 2025
