RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, Atticus Geiger
Abstract
Individual neurons participate in the representation of multiple high-level concepts. To what extent can different interpretability methods successfully disentangle these roles? To help address this question, we introduce RAVEL (Resolving Attribute-Value Entanglements in Language Models), a dataset that enables tightly controlled, quantitative comparisons between a variety of existing interpretability methods. We use the resulting conceptual framework to define the new method of Multi-task Distributed Alignment Search (MDAS), which allows us to find distributed representations satisfying multiple causal criteria. With Llama2-7B as the target language model, MDAS achieves state-of-the-art results on RAVEL, demonstrating the importance of going beyond neuron-level analyses to identify features distributed across activations. We release our benchmark at https://github.com/ explanare/ravel .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- ReFT: Representation Finetuning for Language ModelsZhengxuan Wu, Aryaman Arora, Zheng Wang, Atticus Geiger et al.NeurIPS 2024 · 233 citations
- Do Language Models Use Their Depth Efficiently?Róbert Csordás, Christopher D. Manning, Christopher PottsNeurIPS 2025 · 61 citations
- Automated Interpretability Metrics Do Not Distinguish Trained and Random TransformersThomas Heap, Tim Lawson, Lucy Farnik, Laurence AitchisonICLR 2026 · 32 citations
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez et al.EMNLP 2025 · 10 citations
- Addressing divergent representations from causal interventions on neural networksSatchel Grant, Simon Jerome Han, Alexa R. Tartaglini, Christopher PottsICLR 2026 · 7 citations
Builds on28
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
Related papers
- HyperDAS: Towards Automating Mechanistic Interpretability with HypernetworksJiuding Sun, Jing Huang, Sidharth Baskaran, Karel D'Oosterlinck et al.ICLR 2025
- Interpretability at Scale: Identifying Causal Mechanisms in AlpacaZhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts et al.NeurIPS 2023 · 146 citations
- Wasserstein Distances, Neuronal Entanglement, and SparsityShashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh et al.ICLR 2025
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 11 citations
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 13 citations
