HELM: Hyperbolic Large Language Models via Mixture-of-Curvature Experts
Neil He, Rishabh Anand, Hiren Madhu, Ali Maatouk, Smita Krishnaswamy, Leandros Tassiulas, Menglin Yang, Rex Ying
摘要
Frontier large language models (LLMs) have shown great success in text modeling and generation tasks across domains. However, natural language exhibits inherent semantic hierarchies and nuanced geometric structure, which current LLMs do not capture completely owing to their reliance on Euclidean operations such as dotproducts and norms. Furthermore, recent studies have shown that not respecting the underlying geometry of token embeddings leads to training instabilities and degradation of generative capabilities. These findings suggest that shifting to non-Euclidean geometries can better align language models with the underlying geometry of text. We thus propose to operate fully in Hyperbolic space, known for its expansive, scale-free, and low-distortion properties. To this end, we introduce HELM, a family of HypErbolic Large Language Models, offering a geometric rethinking of the Transformer-based LLM that addresses the representational inflexibility, missing set of necessary operations, and poor scalability of existing hyperbolic LMs. We additionally introduce a Mixture-of-Curvature Experts model, HELM-MICE, where each expert operates in a distinct curvature space to encode more fine-grained geometric structure from text, as well as a dense model, HELM-D. For HELM-MICE, we further develop hyperbolic Multi-Head Latent Attention (HMLA) for efficient, reduced-KV-cache training and inference. For both models, we further develop essential hyperbolic equivalents of rotary positional encodings and root mean square normalization. We are the first to train fully hyperbolic LLMs at billion-parameter scale, and evaluate them on well-known benchmarks such as MMLU and ARC, spanning STEM problem-solving, general knowledge, and commonsense reasoning. Our results show consistent gains from our HELM architectures -up to 4% -over popular Euclidean architectures used in LLaMA and DeepSeek with superior semantic hierarchy modeling capabilities, highlighting the efficacy and enhanced reasoning afforded by hyperbolic geometry in large-scale language model pretraining.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Hyperbolic Busemann Neural NetworksZiheng Chen, Bernhard Schölkopf, Nicu SebeCVPR 2026 · 被引用 4 次
- HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented GenerationHiren Madhu, Ngoc Bui, Ali Maatouk, Leandros Tassiulas 等ICML 2026 · 被引用 1 次
- Hyperbolic Neural Operatorjieyuan pei, Zhuoxuan Li, Wei Li, Haobo Zhang 等ICML 2026
- EEG-Based Multimodal Learning via Hyperbolic Mixture-of-Curvature ExpertsRunhe Zhou, Shanglin Li, Guanxiang Huang, Xinliang Zhou 等ICML 2026
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 被引用 791 次
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu 等ACL 2024 · 被引用 171 次
- Hyperbolic Image-text RepresentationsKaran Desai, Maximilian Nickel, Tanmay Rajpurohit, Justin Johnson 等ICML 2023 · 被引用 137 次
相关 Paper
- Hyperbolic Fine-Tuning for Large Language ModelsMenglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong 等NeurIPS 2025 · 被引用 31 次
- Language Models as Hierarchy EncodersYuan He, Moy Yuan, Jiaoyan Chen, Ian HorrocksNeurIPS 2024 · 被引用 37 次
- Large Language Models Enhanced Hyperbolic Space Recommender SystemsWentao Cheng, Zhida Qin, Zexue Wu, Pengzhan Zhou 等SIGIR 2025 · 被引用 6 次
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language ModelsZelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang 等NeurIPS 2025 · 被引用 5 次
- Hyperbolic Deep Reinforcement LearningEdoardo Cetin, Benjamin Paul Chamberlain, Michael M. Bronstein, Jonathan J. HuntICLR 2023 · 被引用 4 次
