A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations
Bilal Chughtai, Lawrence Chan, Neel Nanda
摘要
Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small neural networks learn to implement group composition. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks of differing architectures trained on various groups, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned -- as well as the order they develop -- are arbitrary.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 被引用 181 次
- Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity TrackingNikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov 等ICLR 2024 · 被引用 113 次
- Successor Heads: Recurring, Interpretable Attention Heads In The WildRhys Gould, Euan Ong, George Ogden, Arthur ConmyICLR 2024 · 被引用 75 次
- Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce GrokkingKaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du 等ICLR 2024 · 被引用 71 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Towards Understanding Grokking: An Effective Theory of Representation LearningZiming Liu, Ouail Kitouni, Niklas Nolte, Eric J. Michaud 等NeurIPS 2022 · 被引用 299 次
- Revisiting Model Stitching to Compare Neural RepresentationsYamini Bansal, Preetum Nakkiran, Boaz BarakNeurIPS 2021 · 被引用 253 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 被引用 183 次
相关 Paper
- Towards a Unified and Verified Understanding of Group-Operation NetworksWilson Wu, Louis Jaburi, Jacob Drori, Jason GrossICLR 2025
- Deep neural networks divide and conquer dihedral multiplicationSihui Wei, Gavin McCracken, Gabriela Moisescu-Pareja, Harley Wiltzer 等ICML 2026
- Grokking Group Multiplication with CosetsDashiell Stander, Qinan Yu, Honglu Fan, Stella BidermanICML 2024 · 被引用 20 次
- Validating Mechanistic Interpretations: An Axiomatic ApproachNils Palumbo, Ravi Mangal, Zifan Wang, Saranya Vijayakumar 等ICML 2025
- Patterning: The Dual of InterpretabilityGeorge Wang, Daniel MurfetICML 2026 · 被引用 5 次
