Real-Time Visual Attribution Streaming in Thinking Model
Seil Kang, Woojung Han, Junhyeok Kim, Jinyeong Kim, Youngeun Kim, Seong Jae Hwang
摘要
We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems from images, their long reasoning traces should be grounded in visual evidence. However, verifying this reliance is challenging: faithful causal methods require costly repeated backward passes or perturbations, while raw attention maps offer instant access, they lack causal validity. To resolve this, we introduce an amortized approach that learns to estimate the causal effects of semantic regions directly from the rich signals encoded in attention features. Across five diverse benchmarks and four thinking models, our approach achieves faithfulness comparable to exhaustive causal methods while enabling visual attribution streaming, where users observe grounding evidence as the model reasons, not after. Our results demonstrate that real-time, faithful attribution in multimodal thinking models is achievable through lightweight learning, not brute-force computation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
相关 Paper
- Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement LearningShuochen Liu, Pengfei Luo, Chao Zhang, Yuhao Chen 等AAAI 2026 · 被引用 2 次
- Multimodal Fact-Level Attribution for Verifiable ReasoningDavid Wan, Han Wang, Ziyang Wang, Elias Stengel-Eskin 等ICML 2026 · 被引用 2 次
- Imagination Helps Visual Reasoning, But Not Yet in Latent SpaceYou Li, Chi Chen, Yanghao Li, Fanhu Zeng 等ICML 2026 · 被引用 6 次
- How Does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse AutoencodingXi Chen, Aske Plaat, Niki van SteinAAAI 2026 · 被引用 9 次
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMsWenbo Pan, Zhichao Liu, Xianlong Wang, Yu Haining 等ICML 2026 · 被引用 3 次
