Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
Yifan Zhang, Wei Bi, Kechi Zhang, Dongming Jin, Jie Fu, Zhi Jin
Abstract
Algorithm extraction aims to synthesize executable programs directly from models trained on algorithmic tasks, enabling de novo recovery of executable mechanisms from weights without relying on human-written target programs. However, applying this paradigm to Transformer is complicated by representation entanglement (e.g., superposition), where features encoded in overlapping directions substantially hinder the recovery of symbolic expressions. We propose the Discrete Transformer, an architecture explicitly designed to bridge the gap between continuous representations and discrete symbolic logic. By injecting discreteness through temperature-annealed sampling, our framework effectively leverages hypothesis testing and symbolic regression to extract human-readable programs. Empirically, the Discrete Transformer achieves performance comparable to the RNN-based MIPS baseline on shared discrete tasks, while broadening extraction to tasks with continuous-valued intermediate computations. Finally, we show that architectural inductive biases provide fine-grained control over synthesized programs, establishing the Discrete Transformer as a controllable testbed for algorithm extraction and Transformer interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Discovering Symbolic Models from Deep Learning with Inductive BiasesMiles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Rui Xu et al.NeurIPS 2020 · 736 citations
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 183 citations
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 162 citations
- Tracr: Compiled Transformers as a Laboratory for InterpretabilityDavid Lindner, János Kramár, Sebastian Farquhar, Matthew Rahtz et al.NeurIPS 2023 · 113 citations
Related papers
- Emergent Discrete Controller Modules for Symbolic Planning in TransformersS. M. Rafiuddin, Muntaha Nujat KhanICLR 2026
- Learning Transformer ProgramsDan Friedman, Alexander Wettig, Danqi ChenNeurIPS 2023 · 59 citations
- UniRPG: Unified Discrete Reasoning over Table and Text as Program GenerationYongwei Zhou, Junwei Bao, Chaoqun Duan, Youzheng Wu et al.EMNLP 2022 · 10 citations
- Algorithmic Capabilities of Random TransformersZiqian Zhong, Jacob AndreasNeurIPS 2024 · 23 citations
- Differentiable Tree Operations Promote Compositional GeneralizationPaul Soulos, Edward J. Hu, Kate McCurdy, Yunmo Chen et al.ICML 2023 · 7 citations
