Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets
Minyoung Hwang, Seokhyun Lee, Changhee Lee
摘要
As deep language models (DLMs) are increasingly deployed in highstakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability. However, achieving this vital level of interpretability is particularly challenging when these DLMs operate as black-box systems (e.g., via APIs), where access to internal model states (e.g., parameters, gradients) is restricted. Despite numerous efforts, existing explanation methods often fail to concurrently satisfy three key desiderata: (i) inference-time efficiency, (ii) black-box compatibility without inducing out-of-distribution behavior, and (iii) comprehensible explanations grounded in the input's linguistic structure. To address these challenges, we propose a method that explains predictions of DLMs by selecting a small, informative subset of input words. We formulate this as an amortized optimization problem, enabling efficient one-shot inference without the need for inputspecific search. Our selection policy is trained via REINFORCEstyle policy gradients, allowing discrete word selection in a fully gradient-free setting. To enhance interpretability and align with human linguistic intuition, we integrate graph-structured knowledge into this selection process, fostering linguistically coherent subsets that result in explanations both highly informative and cognitively meaningful to end-users. We evaluated our method on diverse DLM architectures and multiple real-world datasets. It consistently identifies word subsets with enhanced discriminative power and stronger alignment with linguistically salient cues, outperforming both conventional black-box compatible methods and gradient-based approaches that are given oracle access to the blackbox model's gradients for a more challenging benchmark. Our code is available at here.
• Computing methodologies → Reasoning about belief and knowledge; Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- The Out-of-Distribution Problem in Explainability and Search Methods for Feature Importance ExplanationsPeter Hase, Harry Xie, Mohit BansalNeurIPS 2021 · 被引用 121 次
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause 等NeurIPS 2020 · 被引用 104 次
相关 Paper
- Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionFaeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin ChoiAAAI 2021 · 被引用 46 次
- Generating Hierarchical Explanations on Text Classification via Feature Interaction DetectionHanjie Chen, Guangtao Zheng, Yangfeng JiACL 2020 · 被引用 85 次
- An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare RecordsJoakim Edin, Maria Maistro, Lars Maaløe, Lasse Borgholt 等EMNLP 2024 · 被引用 4 次
- RecExplainer: Aligning Large Language Models for Explaining Recommendation ModelsYuxuan Lei, Jianxun Lian, Jing Yao, Xu Huang 等KDD 2024 · 被引用 18 次
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang 等ICML 2022 · 被引用 343 次
