Can Explanations Be Useful for Calibrating Black Box Models?
Xi Ye, Greg Durrett
摘要
NLP practitioners often want to take existing trained models and apply them to data from new domains. While fine-tuning or few-shot learning can be used to adapt a base model, there is no single recipe for making these techniques work; moreover, one may not have access to the original model weights if it is deployed as a black box. We study how to improve a black box model's performance on a new domain by leveraging explanations of the model's behavior. Our approach first extracts a set of features combining human intuition about the task with model attributions generated by black box interpretation techniques, then uses a simple calibrator, in the form of a classifier, to predict whether the base model was correct or not. We experiment with our method on two tasks, extractive question answering and natural language inference, covering adaptation from several pairs of domains with limited target-domain data. The experimental results across all the domain pairs show that explanations are useful for calibrating these models, boosting accuracy when predictions do not have to be returned on every example. We further show that the calibration model transfers to some extent between tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 被引用 272 次
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas 等ICML 2024 · 被引用 113 次
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng 等EMNLP 2023 · 被引用 91 次
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang 等ICLR 2023 · 被引用 68 次
- ARTiST: Automated Text Simplification for Task Guidance in Augmented RealityGuande Wu, Jing Qian, Sonia Castelo Quispe, Shaoyu Chen 等CHI 2024 · 被引用 20 次
它引用的顶会 Paper5
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 被引用 229 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-BanditsHan Guo, Ramakanth Pasunuru, Mohit BansalAAAI 2020 · 被引用 120 次
- Connecting Attributions and QA Model Behavior on Realistic CounterfactualsXi Ye, Rohan Nair, Greg DurrettEMNLP 2021 · 被引用 13 次
相关 Paper
- CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsAitor Ormazabal, Mikel Artetxe, Eneko AgirreEMNLP 2023 · 被引用 3 次
- Explanation Selection Using Unlabeled Data for Chain-of-Thought PromptingXi Ye, Greg DurrettEMNLP 2023 · 被引用 5 次
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 被引用 52 次
- Learning to Correct for QA Reasoning with Black-box LLMsJaehyung Kim, Dongyoung Kim, Yiming YangEMNLP 2024
- Refining Language Models with Compositional ExplanationsHuihan Yao, Ying Chen, Qinyuan Ye, Xisen Jin 等NeurIPS 2021 · 被引用 39 次
