Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
Zeyuan Allen-Zhu
Abstract
Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tokens), where results are often dominated by noise and randomness. To overcome this, we introduce controlled synthetic pretraining tasks that isolate and evaluate core model capabilities. Within this framework, we discover Canon layers: lightweight architectural components-named after the musical term "canon"-that promote horizontal information flow across neighboring tokens. Canon layers compute weighted sums of nearby token representations and integrate seamlessly into Transformers, linear attention, state-space models, or any sequence architecture. We present 12 key results. This includes how Canon layers enhance reasoning depth (e.g., by 2×), reasoning breadth, knowledge manipulation, etc. They lift weak architectures like NoPE to match RoPE, and linear attention to rival SOTA linear models like Mamba2/GDN-validated both through synthetic tasks and real-world academic-scale pretraining. This synthetic playground offers an economical, principled path to isolate core model capabilities often obscured at academic scales. Equipped with infinite high-quality data, it may even predict how future architectures will behave as training pipelines improve-e.g., through better data curation or RL-based post-training-unlocking deeper reasoning and hierarchical inference.
- Following theory community tradition, we defer the full and future editions of this paper to our project page physics.allen-zhu.com and ssrn.com/abstract=5240330.
The full V1.1 paper underwent NeurIPS 2025 review; due to result density, we recommend consulting the full version for readability. Synthetic GatedDeltaNet (GDN) experiments were added in V2.0; results on 1-8B Canon-layer-pretrained models with real-world data appear in the follow-up Part 4.2 [2]; these were not included in the original NeurIPS 2025 submission, and we reserve the right to submit them elsewhere.
We provide a 3-video tutorial on YouTube: Part 4.1a (methodology & synthetic playground design), Part 4.1b (architecture principles from the playground), and Part 4.2 (when the playground reshapes real-life pretraining).
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e7c282d-47b5-4896-a6c6-ea635c792dd9Cited by top-tier papers14
- Mamba-3: Improved Sequence Modeling using State Space PrinciplesAakash Sunil Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang et al.ICLR 2026 · 96 citations
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri et al.ICML 2026 · 57 citations
- Transformers Provably Learn Chain-of-Thought Reasoning with Length GeneralizationYu Huang, Zixin Wen, Aarti Singh, Yuejie Chi et al.NeurIPS 2025 · 22 citations
- Towards Execution-Grounded Automated AI ResearchChenglei Si, Zitong Yang, Yejin Choi, Emmanuel J Candes et al.ICML 2026 · 12 citations
- Memory Caching: RNNs with Growing MemoryAli Behrouz, Zeman Li, Yuan Deng, Peilin Zhong et al.ICML 2026 · 10 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- Mechanistic Design and Scaling of Hybrid ArchitecturesMichael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy et al.ICML 2024 · 57 citations
- BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural NetworksJong-Hoon Oh, Ryu Iida, Julien Kloetzer, Kentaro TorisawaACL 2021
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 58 citations
- Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural NetworksYixuan Weng, Minjun Zhu, Fei Xia, Bin Li et al.ICLR 2024 · 14 citations
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen et al.NeurIPS 2025 · 39 citations
