Lune

NeurIPS2025顶会

The Fragile Truth of Saliency: Improving LLM Input Attribution via Attention Bias Optimization

Yihua Zhang, Changsheng Wang, Yiwei Chen, Chongyu Fan, Jinghan Jia, Sijia Liu

2025年份
1被引次数

摘要

Input saliency aims to quantify the influence of input tokens on the output of large language models (LLMs), which has been widely used for prompt engineering, model interpretability, and behavior attribution. Despite the proliferation of saliency techniques, the field lacks a standardized and rigorous evaluation protocol. In this work, we introduce a stress-testing framework inspired by the needle-in-a-haystack (NIAH) setting to systematically assess the reliability of seven popular input saliency methods. Our evaluation reveals a surprising and critical flaw: existing methods consistently assign non-trivial importance to irrelevant context, and this attribution error worsens as input length increases. To address this issue, we propose a novel saliency method based on Attention Bias Optimization (ABO), which explicitly optimizes the attention bias associated with each input token to quantify its causal impact on target token generation. ABO robustly outperforms existing methods by 10 ∼ 30% in saliency accuracy across diverse NIAH tasks, maintains effectiveness up to 10K-token prompts, and enables practical applications including zero-shot detoxification, sentiment steering, and reasoning-error correction. Our findings highlight the limitations of prevalent attribution methods and establish ABO as a principled alternative for accurate token attribution.

• We introduce a stress-testing framework for LLM saliency based on the needle-in-a-haystack task, enabling fine-grained diagnosis under long-context inputs.

• We reveal systemic reliability failures in six widely-used attribution methods and show that misattribution to irrelevant tokens can exceed over 90 % in 10 K-token prompts, quantifying the brittleness of existing saliency tools.

• We propose Attention Bias Optimization (ABO), a principled, optimization-based technique that delivers token-level saliency scores with superior causal fidelity and scalability to long-contexts.

• We demonstrate ABO's broad practical utility across sentiment control, toxic prompt detoxification, and LLM error correction, underscoring its value for both research and real-world deployment.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper33

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖