Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and Recoverability
Advaith Malladi, Shashank Srivastava
Abstract
Natural-language explanations are widely used to interpret machine learning models, yet many prioritize human plausibility over accurately reflecting or predicting model behavior. Prior approaches often rely on human-written rationales, producing post-hoc explanations that neither align with the model's decision function nor generalize. We introduce OPEX , a natural-language explanation model that directly optimizes for behavioral faithfulness: the ability of an explanation to reflect and predict a model's observable input-output behavior. OPEX is trained using reinforcement learning with Group Relative Policy Optimization (GRPO), optimizing two complementary metrics: recoverability, which measures whether explanations recover model predictions on seen examples, and simulatability, which measures prediction of model behavior on unseen inputs. Across structured and text-based tasks, OPEX achieves high simulatability (∼0.85) and recoverability (∼0.99), outperforming GPT-4o, LLaMA-3.3-70B, MaNtLE, Chain-of-Thought (CoT)-based models, and human-written explanations, despite using an 8B-parameter backbone. Human user studies show a 15% improvement in classification accuracy over competent baselines. Link to Code and OPEX weights: github.com/advaithmall/OPeX
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- CLUES: A Benchmark for Learning Classifiers using Natural Language ExplanationsRakesh R. Menon, Sayan Ghosh, Shashank SrivastavaACL 2022 · 13 citations
Related papers
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionAaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan et al.ICML 2022 · 48 citations
- GRACE: Generative Representation Learning via Contrastive Policy OptimizationJiashuo Sun, Shixuan Liu, Zhaochen Su, Xianrui Zhong et al.ICLR 2026 · 7 citations
- MaNtLE: Model-agnostic Natural Language ExplainerRakesh R. Menon, Kerem Zaman, Shashank SrivastavaEMNLP 2023 · 1 citation
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceBar Alon, Itamar Zimerman, Lior WolfACL 2026
