Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
Yifan Zhang, Wei Bi, Kechi Zhang, Dongming Jin, Jie Fu, Zhi Jin
摘要
Algorithm extraction aims to synthesize executable programs directly from models trained on algorithmic tasks, enabling de novo recovery of executable mechanisms from weights without relying on human-written target programs. However, applying this paradigm to Transformer is complicated by representation entanglement (e.g., superposition), where features encoded in overlapping directions substantially hinder the recovery of symbolic expressions. We propose the Discrete Transformer, an architecture explicitly designed to bridge the gap between continuous representations and discrete symbolic logic. By injecting discreteness through temperature-annealed sampling, our framework effectively leverages hypothesis testing and symbolic regression to extract human-readable programs. Empirically, the Discrete Transformer achieves performance comparable to the RNN-based MIPS baseline on shared discrete tasks, while broadening extraction to tasks with continuous-valued intermediate computations. Finally, we show that architectural inductive biases provide fine-grained control over synthesized programs, establishing the Discrete Transformer as a controllable testbed for algorithm extraction and Transformer interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu 等NeurIPS 2020 · 被引用 736 次
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 被引用 183 次
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 被引用 162 次
- Tracr: Compiled Transformers as a Laboratory for InterpretabilityDavid Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz 等NeurIPS 2023 · 被引用 113 次
相关 Paper
- Emergent Discrete Controller Modules for Symbolic Planning in TransformersS. M. Rafiuddin, Muntaha Nujat KhanICLR 2026
- Learning Transformer ProgramsDan Friedman, Alexander Wettig, Danqi ChenNeurIPS 2023 · 被引用 59 次
- UniRPG: Unified Discrete Reasoning over Table and Text as Program GenerationYongwei Zhou, Junwei Bao, Chaoqun Duan, Youzheng Wu 等EMNLP 2022 · 被引用 10 次
- Algorithmic Capabilities of Random TransformersZiqian Zhong, Jacob AndreasNeurIPS 2024 · 被引用 23 次
- Differentiable Tree Operations Promote Compositional GeneralizationPaul Soulos, Edward J. Hu, Kate McCurdy, Yunmo Chen 等ICML 2023 · 被引用 7 次
