Multi-Level Explanations for Generative Language Models
Lucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt, Ronny Luss, Amit Dhurandhar, Manish Nagireddy, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Werner Geyer, Soumya Ghosh
摘要
Despite the increasing use of large language models (LLMs) for context-grounded tasks like summarization and question-answering, understanding what makes an LLM produce a certain response is challenging. We propose Multi-Level Explanations for Generative Language Models (MExGen), a technique to provide explanations for context-grounded text generation. MExGen assigns scores to parts of the context to quantify their influence on the model's output. It extends attribution methods like LIME and SHAP to LLMs used in context-grounded tasks where (1) inference cost is high, (2) input text is long, and (3) the output is text. We conduct a systematic evaluation, both automated and human, of perturbation-based attribution methods for summarization and question answering. The results show that our framework can provide more faithful explanations of generated output than available alternatives, including LLM self-explanations. We open-source code for MExGen as part of the ICX360 toolkit: https://github$.$com/IBM/ICX360.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning to Understand: Identifying Interactions via the Möbius TransformJustin Singh Kang, Yigit Efe Erginbas, Landon Butler, Ramtin Pedarsani 等NeurIPS 2024 · 被引用 17 次
- ReX: A Framework for Incorporating Temporal Information in Model-Agnostic Local Explanation TechniquesJunhao Liu, Xin ZhangAAAI 2025 · 被引用 6 次
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 被引用 4 次
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy ModelsJunhao Liu, Haonan Yu, Zhenyu Yan, Xin ZhangACL 2026 · 被引用 2 次
- EnsembleSHAP: Faithful and Certifiably Robust Attribution for Random Subspace MethodYanting Wang, Jinyuan JiaICLR 2026
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- JoPA: Explaining Large Language Model's Generation via Joint Prompt AttributionYurui Chang, Bochuan Cao, Yujia Wang, Jinghui Chen 等ACL 2025
- LAQuer: Localized Attribution Queries in Content-grounded GenerationEran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin 等ACL 2025
- On Synthesizing Data for Context Attribution in Question AnsweringGorjan Radevski, Kiril Gashteovski, Shahbaz Syed, Christopher Malon 等ACL 2025 · 被引用 1 次
- TracLLM: A Generic Framework for Attributing Long Context LLMsYanting Wang, Wei Zou, Runpeng Geng, Jinyuan JiaUSENIX Security 2025
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
