The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization
Yihua Zhang, Changsheng Wang, Yiwei Chen, Chongyu Fan, Jinghan Jia, Sijia Liu
摘要
Input saliency aims to quantify the influence of input tokens on the output of large language models (LLMs), which has been widely used for prompt engineering, model interpretability, and behavior attribution. Despite the proliferation of saliency techniques, the field lacks a standardized and rigorous evaluation protocol. In this work, we introduce a stress-testing framework inspired by the needle-in-a-haystack (NIAH) setting to systematically assess the reliability of seven popular input saliency methods. Our evaluation reveals a surprising and critical flaw: existing methods consistently assign non-trivial importance to irrelevant context, and this attribution error worsens as input length increases. To address this issue, we propose a novel saliency method based on Attention Bias Optimization (ABO), which explicitly optimizes the attention bias associated with each input token to quantify its causal impact on target token generation. ABO robustly outperforms existing methods by 10 ∼ 30% in saliency accuracy across diverse NIAH tasks, maintains effectiveness up to 10K-token prompts, and enables practical applications including zero-shot detoxification, sentiment steering, and reasoning-error correction. Our findings highlight the limitations of prevalent attribution methods and establish ABO as a principled alternative for accurate token attribution.
• We introduce a stress-testing framework for LLM saliency based on the needle-in-a-haystack task, enabling fine-grained diagnosis under long-context inputs.
• We reveal systemic reliability failures in six widely-used attribution methods and show that misattribution to irrelevant tokens can exceed over 90 % in 10 K-token prompts, quantifying the brittleness of existing saliency tools.
• We propose Attention Bias Optimization (ABO), a principled, optimization-based technique that delivers token-level saliency scores with superior causal fidelity and scalability to long-contexts.
• We demonstrate ABO's broad practical utility across sentiment control, toxic prompt detoxification, and LLM error correction, underscoring its value for both research and real-world deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- Unveiling and Manipulating Prompt Influence in Large Language ModelsZijian Feng, Hanzhang Zhou, Zixiao Zhu, Junlang Qian 等ICLR 2024 · 被引用 9 次
- Attention Sinks as Internal Signals for Hallucination Detection in Large Language ModelsJakub Binkowski, Kamil Adamczewski, Tomasz KajdanowiczICML 2026
- Attention Speaks Volumes: Localizing and Mitigating Bias in Language ModelsRishabh Adiga, Besmira Nushi, Varun ChandrasekaranACL 2025
- GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language ModelsZhibo Zhang, Wuxia Bai, Yuxi Li, Mark Huasong Meng 等ASE 2024 · 被引用 2 次
- NoLiMa: Long-Context Evaluation Beyond Literal MatchingAli Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui 等ICML 2025
