Can Explanations Be Useful for Calibrating Black Box Models?
Xi Ye, Greg Durrett
Abstract
NLP practitioners often want to take existing trained models and apply them to data from new domains. While fine-tuning or few-shot learning can be used to adapt a base model, there is no single recipe for making these techniques work; moreover, one may not have access to the original model weights if it is deployed as a black box. We study how to improve a black box model's performance on a new domain by leveraging explanations of the model's behavior. Our approach first extracts a set of features combining human intuition about the task with model attributions generated by black box interpretation techniques, then uses a simple calibrator, in the form of a classifier, to predict whether the base model was correct or not. We experiment with our method on two tasks, extractive question answering and natural language inference, covering adaptation from several pairs of domains with limited target-domain data. The experimental results across all the domain pairs show that explanations are useful for calibrating these models, boosting accuracy when predictions do not have to be returned on every example. We further show that the calibration model transfers to some extent between tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5683cc2e-9f92-406d-a3e9-63fa134af587Cited by top-tier papers9
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
- Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingBairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas et al.ICML 2024 · 113 citations
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng et al.EMNLP 2023 · 91 citations
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang et al.ICLR 2023 · 68 citations
- ARTiST: Automated Text Simplification for Task Guidance in Augmented RealityGuande Wu, Jing Qian, Sonia Castelo Quispe, Shaoyu Chen et al.CHI 2024 · 20 citations
Builds on5
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-BanditsHan Guo, Ramakanth Pasunuru, Mohit BansalAAAI 2020 · 120 citations
- Connecting Attributions and QA Model Behavior on Realistic CounterfactualsXi Ye, Rohan Nair, Greg DurrettEMNLP 2021 · 13 citations
Related papers
- CombLM: Adapting Black-Box Language Models through Small Fine-Tuned ModelsAitor Ormazabal, Mikel Artetxe, Eneko AgirreEMNLP 2023 · 3 citations
- Explanation Selection Using Unlabeled Data for Chain-of-Thought PromptingXi Ye, Greg DurrettEMNLP 2023 · 5 citations
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 52 citations
- Learning to Correct for QA Reasoning with Black-box LLMsJaehyung Kim, Dongyoung Kim, Yiming YangEMNLP 2024
- Refining Language Models with Compositional ExplanationsHuihan Yao, Ying Chen, Qinyuan Ye, Xisen Jin et al.NeurIPS 2021 · 39 citations
