mHC: Manifold-Constrained Hyper-Connections
Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Mingyu Xu, Kuai Yu, Liang Zhao, Shangyan Zhou
摘要
Recently, studies exemplified by Hyper-Connections (HC) have extended the ubiquitous residual connection paradigm established over the past decade by expanding the residual stream width and diversifying connectivity patterns. While yielding substantial performance gains, this diversification fundamentally compromises the identity mapping property intrinsic to the residual connection, which causes severe training instability and restricted scalability, and additionally incurs notable memory access overhead. To address these challenges, we propose Manifold-Constrained Hyper-Connections (mHC), a general framework that projects the residual connection space of HC onto a specific manifold to restore the identity mapping property, while incorporating rigorous infrastructure optimization to ensure efficiency. Empirical experiments demonstrate that mHC is effective for training at scale, offering tangible performance improvements and superior scalability. We anticipate that mHC, as a flexible and practical extension of HC, will contribute to a deeper understanding of topological architecture design and suggest promising directions for the evolution of foundational models. (a) Residual Connection (b) Hyper-Connections (HC) (c) Manifold-Constrained HC (mHC) Layer ℱ x ! x !"# Res Mapping ℋ ! $%&
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen 等ACL 2026 · 被引用 57 次
- Controlled LLM Training on Spectral SphereTian Xie, Haoming Luo, Haoyu Tang, Hu Yiwen 等ICML 2026 · 被引用 22 次
- SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-NormTianyu Li, Dongchen Han, Zixuan Cao, Haofeng Huang 等ICML 2026 · 被引用 7 次
- When Does Sparsity Mitigate the Curse of Depth in LLMsYao Yao, Xinyuan Song, Sebastian Pokutta, Max Zimmer 等ICML 2026 · 被引用 5 次
- Why Steering Works: Toward a Unified View of Language Model Parameter DynamicsZiwen Xu, Chenyan Wu, Hengyu Sun, Haiwen Hong 等ACL 2026 · 被引用 4 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
相关 Paper
- Hyper-ConnectionsDefa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng 等ICLR 2025
- KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual MatricesWuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo MandicICML 2026 · 被引用 8 次
- Learning in Compact Spaces with Approximately Normalized TransformerJörg K. H. Franke, Urs Spiegelhalter, Marianna Nezhurina, Jenia Jitsev 等NeurIPS 2025 · 被引用 5 次
- Residual Hyperbolic Graph Convolution NetworksYangkai Xue, Jindou Dai, Zhipeng Lu, Yuwei Wu 等AAAI 2024 · 被引用 10 次
- LeHDC: learning-based hyperdimensional computing classifierShijin Duan, Yejia Liu, Shaolei Ren, Xiaolin XuDAC 2022 · 被引用 38 次
