Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
Zeyuan Allen-Zhu
摘要
Understanding architectural differences in language models is challenging, especially at academic-scale pretraining (e.g., 1.3B parameters, 100B tokens), where results are often dominated by noise and randomness. To overcome this, we introduce controlled synthetic pretraining tasks that isolate and evaluate core model capabilities. Within this framework, we discover Canon layers: lightweight architectural components-named after the musical term "canon"-that promote horizontal information flow across neighboring tokens. Canon layers compute weighted sums of nearby token representations and integrate seamlessly into Transformers, linear attention, state-space models, or any sequence architecture. We present 12 key results. This includes how Canon layers enhance reasoning depth (e.g., by 2×), reasoning breadth, knowledge manipulation, etc. They lift weak architectures like NoPE to match RoPE, and linear attention to rival SOTA linear models like Mamba2/GDN-validated both through synthetic tasks and real-world academic-scale pretraining. This synthetic playground offers an economical, principled path to isolate core model capabilities often obscured at academic scales. Equipped with infinite high-quality data, it may even predict how future architectures will behave as training pipelines improve-e.g., through better data curation or RL-based post-training-unlocking deeper reasoning and hierarchical inference.
- Following theory community tradition, we defer the full and future editions of this paper to our project page physics.allen-zhu.com and ssrn.com/abstract=5240330.
The full V1.1 paper underwent NeurIPS 2025 review; due to result density, we recommend consulting the full version for readability. Synthetic GatedDeltaNet (GDN) experiments were added in V2.0; results on 1-8B Canon-layer-pretrained models with real-world data appear in the follow-up Part 4.2 [2]; these were not included in the original NeurIPS 2025 submission, and we reserve the right to submit them elsewhere.
We provide a 3-video tutorial on YouTube: Part 4.1a (methodology & synthetic playground design), Part 4.1b (architecture principles from the playground), and Part 4.2 (when the playground reshapes real-life pretraining).
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Mamba-3: Improved Sequence Modeling using State Space PrinciplesAakash Sunil Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang 等ICLR 2026 · 被引用 96 次
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri 等ICML 2026 · 被引用 57 次
- Transformers Provably Learn Chain-of-Thought Reasoning with Length GeneralizationYu Huang, Zixin Wen, Aarti Singh, Yuejie Chi 等NeurIPS 2025 · 被引用 22 次
- Towards Execution-Grounded Automated AI ResearchChenglei Si, Zitong Yang, Yejin Choi, Emmanuel J Candes 等ICML 2026 · 被引用 12 次
- Memory Caching: RNNs with Growing MemoryAli Behrouz, Zeman Li, Yuan Deng, Peilin Zhong 等ICML 2026 · 被引用 10 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- Mechanistic Design and Scaling of Hybrid ArchitecturesMichael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy 等ICML 2024 · 被引用 57 次
- BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural NetworksJong-Hoon Oh, Ryu Iida, Julien Kloetzer, Kentaro TorisawaACL 2021
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 被引用 58 次
- Mastering Symbolic Operations: Augmenting Language Models with Compiled Neural NetworksYixuan Weng, Minjun Zhu, Fei Xia, Bin Li 等ICLR 2024 · 被引用 14 次
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen 等NeurIPS 2025 · 被引用 39 次
