mHC: Manifold-Constrained Hyper-Connections
Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Mingyu Xu, Kuai Yu, Liang Zhao, Shangyan Zhou
Abstract
Recently, studies exemplified by Hyper-Connections (HC) have extended the ubiquitous residual connection paradigm established over the past decade by expanding the residual stream width and diversifying connectivity patterns. While yielding substantial performance gains, this diversification fundamentally compromises the identity mapping property intrinsic to the residual connection, which causes severe training instability and restricted scalability, and additionally incurs notable memory access overhead. To address these challenges, we propose Manifold-Constrained Hyper-Connections (mHC), a general framework that projects the residual connection space of HC onto a specific manifold to restore the identity mapping property, while incorporating rigorous infrastructure optimization to ensure efficiency. Empirical experiments demonstrate that mHC is effective for training at scale, offering tangible performance improvements and superior scalability. We anticipate that mHC, as a flexible and practical extension of HC, will contribute to a deeper understanding of topological architecture design and suggest promising directions for the evolution of foundational models. (a) Residual Connection (b) Hyper-Connections (HC) (c) Manifold-Constrained HC (mHC) Layer ℱ x ! x !"# Res Mapping ℋ ! $%&
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee7ae91d-f53f-4dc2-b92b-bb3fd161f708Cited by top-tier papers11
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen et al.ACL 2026 · 57 citations
- Controlled LLM Training on Spectral SphereTian Xie, Haoming Luo, Haoyu Tang, Hu Yiwen et al.ICML 2026 · 22 citations
- SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-NormTianyu Li, Dongchen Han, Zixuan Cao, Haofeng Huang et al.ICML 2026 · 7 citations
- When Does Sparsity Mitigate the Curse of Depth in LLMsYao Yao, Xinyuan Song, Sebastian Pokutta, Max Zimmer et al.ICML 2026 · 5 citations
- Why Steering Works: Toward a Unified View of Language Model Parameter DynamicsZiwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong et al.ACL 2026 · 4 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng et al.ICLR 2025
- KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual MatricesWuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo MandicICML 2026 · 8 citations
- Learning in Compact Spaces with Approximately Normalized TransformerJörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev et al.NeurIPS 2025 · 5 citations
- Residual Hyperbolic Graph Convolution NetworksYangkai Xue, Jindou Dai, Zhipeng Lu, Yuwei Wu et al.AAAI 2024 · 10 citations
- LeHDC: learning-based hyperdimensional computing classifierShijin Duan, Yejia Liu, Shaolei Ren, Xiaolin XuDAC 2022 · 38 citations
