Circuit Stability Characterizes Language Model Generalization
Alan Sun
摘要
Extensively evaluating the capabilities of (large) language models is difficult. Rapid development of state-of-the-art models induce benchmark saturation, while creating more challenging datasets is labor-intensive. Inspired by the recent developments in mechanistic interpretability, we introduce circuit stability as a new way to assess model performance. Circuit stability refers to a model's ability to apply a consistent reasoning process-its circuitacross various inputs. We mathematically formalize circuit stability and circuit equivalence. Then, through three case studies, we empirically show that circuit stability and the lack thereof can characterize and predict different aspects of generalization. Our proposed methods offer a step towards rigorously relating the generality of models to their interpretability 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper30
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
相关 Paper
- Towards Global-level Mechanistic Interpretability: A Perspective of Modular Circuits of Large Language ModelsYinhan He, Wendy Zheng, Yushun Dong, Yaochen Zhu 等ICML 2025
- Circuit Component Reuse Across Tasks in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickICLR 2024 · 被引用 108 次
- Circuit Compositions: Exploring Modular Structures in Transformer-Based Language ModelsPhilipp Mondorf, Sondre Wold, Barbara PlankACL 2025 · 被引用 5 次
- Mechanistic Interpretability as Statistical Estimation: A Variance AnalysisMaxime Méloux, François Portet, Maxime PeyrardICML 2026 · 被引用 13 次
- Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster InferenceJorge García-Carrasco, Alejandro Maté, Juan TrujilloAAAI 2025 · 被引用 3 次
