Explaining Differences Between Model Pairs in Natural Language through Sample Learning
Advaith Malladi, Rakesh R. Menon, Yuvraj Jain, Shashank Srivastava
摘要
With the growing adoption of machine learning models in critical domains, techniques for explaining differences between models have become essential for trust, debugging, and informed deployment. Previous approaches address this by identifying input transformations that cause divergent predictions (Shah et al., 2022) or by learning joint surrogate models to align and contrast behaviors (Haldar et al., 2023) . These methods often require access to training data and do not produce natural language explanations. In this paper, we introduce SLED , a framework that generates faithful natural language explanations of when and how two ML models converge or diverge in their predictions. SLED first uses gradient-based optimization to synthesize input samples that highlight divergence and convergence patterns, and then leverages a large language model (LLM) to generate explanations grounded in these synthetic samples. Across both text-based (3 tasks, 7 models) and structured (10 tasks, 4 models) classification tasks, we show that SLED explanations are 18-24% more faithful than the strongest baselines. User studies also indicate that SLED explanations achieve a real-world simulatability of 63.5%. Importantly, SLED requires minimal access to training data and generalizes well to real-world samples, enabling transparent model comparison. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 被引用 216 次
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 被引用 102 次
- Faithful Explanations of Black-box NLP Models Using LLM-generated CounterfactualsYair Ori Gat, Nitay Calderon, Amir Feder, Alexander Chapanin 等ICLR 2024 · 被引用 55 次
- ModelDiff: A Framework for Comparing Learning AlgorithmsHarshay Shah, Sung Min Park, Andrew Ilyas, Aleksander MadryICML 2023 · 被引用 36 次
相关 Paper
- MaNtLE: Model-agnostic Natural Language ExplainerRakesh R. Menon, Kerem Zaman, Shashank SrivastavaEMNLP 2023 · 被引用 1 次
- Explain the Synth: Interpretable Evaluation of LLM Data SynthesisYue Yang, Fan Yang, Yu Bai, Hao WangACL 2026
- Making Sense of LLM Decisions: A Prototype-based Framework for Explainable ClassificationBowen Wei, Mehrdad Fazli, Ziwei ZhuAAAI 2026
- Constraint-Driven Explanations for Black-Box ML ModelsAditya A. Shrotri, Nina Narodytska, Alexey Ignatiev, Kuldeep S. Meel 等AAAI 2022 · 被引用 25 次
- Explaining with Contrastive Phrasal Highlighting: A Case Study in Assisting Humans to Detect Translation DifferencesEleftheria Briakou, Navita Goyal, Marine CarpuatEMNLP 2023
