Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
Angie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz, Arvind Satyanarayan, Fred Hohman
摘要
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions can be vague, inconsistent, and require manual relabeling. In response, we introduce semantic regexes, structured language descriptions of LLM features. By combining primitives that capture linguistic and semantic patterns with modifiers for contextualization, composition, and quantification, semantic regexes produce precise and expressive feature descriptions. Across quantitative benchmarks and qualitative analyses, semantic regexes match the accuracy of natural language while yielding more concise and consistent feature descriptions. Their inherent structure affords new types of analyses, including quantifying feature complexity across layers, scaling automated interpretability from insights into individual features to model-wide patterns. Finally, in user studies, we find that semantic regexes help people build accurate mental models of LLM features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 被引用 222 次
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar 等NeurIPS 2025 · 被引用 168 次
- Natural Language Descriptions of Deep Visual FeaturesEvan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili 等ICLR 2022 · 被引用 160 次
相关 Paper
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer 等NeurIPS 2025 · 被引用 12 次
- Enhancing Automated Interpretability with Output-Centric Feature DescriptionsYoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger 等ACL 2025
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsHaoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li 等IEEE VIS 2025 · 被引用 3 次
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha 等ICML 2026 · 被引用 6 次
