Composing Linear Layers from Irreducibles
Travis Pence, Daisuke Yamada, Vikas Singh
摘要
Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood. We investigate this compositional structure in linear layers by asking: can we identify/synthesize linear transformations from a minimal set of geometric primitives? Using Clifford algebra, we show that linear layers can be expressed as compositions of bivectors -- geometric objects encoding oriented planes -- and introduce a differentiable algorithm that decomposes them into products of rotors. This construction uses only O(log^2 d) parameters, versus O(d^2) required by dense matrices. Applied to the key, query, and value projections in LLM attention layers, our rotor-based layers match the performance of strong baselines such as block-Hadamard and low-rank approximations. Our findings provide an algebraic perspective on how these geometric primitives can compose into higher-level functions within deep models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Thinking Like TransformersGail Weiss, Yoav Goldberg, Eran YahavICML 2021 · 被引用 183 次
- Compositional Generalization via Neural-Symbolic Stack MachinesXinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song 等NeurIPS 2020 · 被引用 112 次
- Clifford Group Equivariant Neural NetworksDavid Ruhe, Johannes Brandstetter, Patrick ForréNeurIPS 2023 · 被引用 85 次
相关 Paper
- Geometric Algebra TransformerJohann Brehmer, Pim de Haan, Sönke Behrends, Taco S. CohenNeurIPS 2023 · 被引用 81 次
- Geometric Clifford Algebra NetworksDavid Ruhe, Jayesh K. Gupta, Steven De Keninck, Max Welling 等ICML 2023 · 被引用 58 次
- GLGENN: A Novel Parameter-Light Equivariant Neural Networks Architecture Based on Clifford Geometric AlgebrasEkaterina Filimoshina, Dmitry ShirokovICML 2025
- Clifford Group Equivariant Simplicial Message Passing NetworksCong Liu, David Ruhe, Floor Eijkelboom, Patrick ForréICLR 2024 · 被引用 20 次
- Clifford Neural Layers for PDE ModelingJohannes Brandstetter, Rianne van den Berg, Max Welling, Jayesh K. GuptaICLR 2023 · 被引用 21 次
