Grokking Group Multiplication with Cosets
Dashiell Stander, Qinan Yu, Honglu Fan, Stella Biderman
摘要
The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep neural networks, but all have substantial limitations. Algorithmic tasks have proven to be a fruitful test ground for interpreting a neural network end-to-end. Building on previous work, we completely reverse engineer fully connected one-hidden layer networks that have ``grokked'' the arithmetic of the permutation groups and . The models discover the true subgroup structure of the full group and converge on neural circuits that decompose the group arithmetic using the permutation group's subgroups. We relate how we reverse engineered the model's mechanisms and confirmed our theory was a faithful description of the circuit's functionality. We also draw attention to current challenges in conducting interpretability research by comparing our work to Chughtai et al. [4] which alleges to find a different algorithm for this same problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Hypothesis Testing the Circuit Hypothesis in LLMsClaudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng 等NeurIPS 2024 · 被引用 36 次
- Uncovering a Universal Abstract Algorithm for Modular Addition in Neural NetworksGavin McCracken, Gabriela Moisescu-Pareja, Vincent Létourneau, Doina Precup 等NeurIPS 2025 · 被引用 14 次
- Sequential Group Composition: A Window into the Mechanics of Deep LearningGiovanni Luca Marchetti, Daniel Kunin, Adele Myers, Francisco Acosta 等ICML 2026 · 被引用 8 次
- On The Geometry and Topology of Representations: the Manifolds of Modular AdditionGabriela Moisescu-Pareja, Gavin McCracken, Harley Wiltzer, Colin Daniels 等ICLR 2026 · 被引用 3 次
- In-Context AlgebraEric Todd, Jannik Brinkmann, Rohit Gandikota, David BauICLR 2026 · 被引用 3 次
它引用的顶会 Paper16
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language ModelsPeter Hase, Mohit Bansal, Been Kim, Asma GhandehariounNeurIPS 2023 · 被引用 307 次
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 被引用 181 次
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
相关 Paper
- Towards a Unified and Verified Understanding of Group-Operation NetworksWilson Wu, Louis Jaburi, Jacob Drori, Jason GrossICLR 2025
- Deep neural networks divide and conquer dihedral multiplicationSihui Wei, Gavin McCracken, Gabriela Moisescu-Pareja, Harley Wiltzer 等ICML 2026
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- Learning to Add, Multiply, and Execute Algorithmic Instructions Exactly with Neural NetworksArtur Back de Luca, George Giapitzakis, Kimon FountoulakisNeurIPS 2025 · 被引用 4 次
- On the Symmetries of Deep Learning Models and their Internal RepresentationsCharles Godfrey, Davis Brown, Tegan Emerson, Henry KvingeNeurIPS 2022 · 被引用 78 次
