Hyper-SET: Designing Transformers via Hyperspherical Energy Minimization
Yunzhe Hu, Difan Zou, Dong Xu
摘要
Transformer-based models have achieved remarkable success, but their core components, Transformer layers, are largely heuristics-driven and engineered from the bottom up, calling for a prototypical model with high interpretability and practical competence. To this end, we conceptualize a principled, top-down approach grounded in energy-based interpretation. Specifically, we formalize token dynamics as a joint maximum likelihood estimation on the hypersphere, featuring two properties: semantic alignment in the high-dimensional space and distributional uniformity in the low-dimensional space. By quantifying them with extended Hopfield energy functions, we instantiate this idea as a constrained energy minimization problem, which enables designs of symmetric attention and feedforward modules with RMS normalization. We further present Hyper-Spherical Energy Transformer (Hyper-SET), a recurrent-depth alternative to vanilla Transformers naturally emerging from iterative energy optimization on the hypersphere. With shared parameters across layers, Hyper-SET can scale to arbitrary depth with fewer parameters. Theoretically grounded and compact, it achieves competitive or superior performance across diverse tasks, including Sudoku solving, image classification, and masked image modeling. We also design novel variations under the proposed general principle, such as linear attention and gated feedforward layer, and showcase its scalability with depth-wise LoRA. Our results highlight Hyper-SET as a step toward interpretable and principled Transformer design. Code is available at https://github.com/huyunzhe/hyper-set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from JacobiansAkiyoshi Tomihari, Ryo KarakidaNeurIPS 2025 · 被引用 5 次
- C-Voting: Confidence-Based Test-Time Voting without Explicit Energy FunctionsKenji Kubo, Shunsuke Kamiya, Masanori Koyama, Kohei Hayashi 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper45
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Transformers from an Optimization PerspectiveYongyi Yang, Zengfeng Huang, David P. WipfNeurIPS 2022 · 被引用 44 次
- nGPT: Normalized Transformer with Representation Learning on the HypersphereIlya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris GinsburgICLR 2025
- ETHER: Efficient Finetuning of Large-Scale Models with Hyperplane ReflectionsMassimo Bini, Karsten Roth, Zeynep Akata, Anna KhorevaICML 2024 · 被引用 11 次
- Energy-Regularized Sequential Model Editing on HyperspheresQingyuan Liu, Jia-Chen Gu, Yunzhi Yao, Hong Wang 等ICLR 2026 · 被引用 1 次
- QUEST: A robust attention formulation using query-modulated spherical attentionHariprasath Govindarajan, Per Sidén, Jacob Roll, Fredrik LindstenICLR 2026 · 被引用 1 次
