Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
Jing Huang, Junyi Tao, Thomas Icard, Diyi Yang, Christopher Potts
Abstract
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-ofdistribution examples? In this work, we provide a positive answer to this question. Through a diverse set of language modeling tasks-including symbol manipulation, knowledge retrieval, and instruction following-we show that the most robust features for correctness prediction are those that play a distinctive causal role in the model's behavior. Specifically, we propose two methods that leverage causal mechanisms to predict the correctness of model outputs: counterfactual simulation (checking whether key causal variables are realized) and value probing (using the values of those variables to make predictions). Both achieve high AUC-ROC in distribution and outperform methods that rely on causal-agnostic features in out-of-distribution settings, where predicting model behaviors is more crucial. Our work thus highlights a novel and significant application for internal causal analysis of language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49523918-4f51-400f-8198-fe1f2f61a035Builds on21
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
Related papers
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 5 citations
- Predicting Fine-Tuning Performance with ProbingZining Zhu, Soroosh Shahtalebi, Frank RudziczEMNLP 2022 · 6 citations
- CausalGym: Benchmarking causal interpretability methods on linguistic tasksAryaman Arora, Dan Jurafsky, Christopher PottsACL 2024 · 3 citations
- Causal Interventions Reveal Shared Structure Across English Filler-Gap ConstructionsSasha Boguraev, Christopher Potts, Kyle MahowaldEMNLP 2025
- Predicting the Performance of Black-box Language Models with Follow-up QueriesDylan Sam, Marc Finzi, Zico KolterNeurIPS 2025 · 10 citations
