Hyperbolic Fine-Tuning for Large Language Models
Menglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong, Jiahong Liu, Irwin King, Rex Ying
摘要
Large language models (LLMs) have demonstrated remarkable performance across various tasks. However, it remains an open question whether the default Euclidean space is the most suitable choice for LLMs. In this study, we investigate the geometric characteristics of LLMs, focusing specifically on tokens and their embeddings. Our findings reveal that token frequency follows a power-law distribution, where high-frequency tokens (e.g., "the," "that") constitute the minority, while low-frequency tokens (e.g., "apple," "dog") constitute the majority. Furthermore, high-frequency tokens cluster near the origin, whereas low-frequency tokens are positioned farther away in the embedding space. Additionally, token embeddings exhibit hyperbolic characteristics, indicating a latent tree-like structure within the embedding space. Motivated by these observations, we propose HypLoRA, an efficient fine-tuning approach that operates in hyperbolic space to exploit these underlying hierarchical structures better. HypLoRA performs low-rank adaptation directly in hyperbolic space, thereby preserving hyperbolic modeling capabilities throughout the fine-tuning process. Extensive experiments across various base models and reasoning benchmarks, specifically arithmetic and commonsense reasoning tasks, demonstrate that HypLoRA substantially improves LLM performance.
Recent advancements suggest that non-Euclidean geometries, particularly hyperbolic spaces [11,14,15,16,17,18,19,20,21,22,23], offer promising alternatives for modeling hierarchical data. Hyperbolic space, distinguished by its negative curvature, is especially well-suited for representing tree-like hierarchical data due to its exponential volume growth and geometric prior. This geometric property makes hyperbolic space particularly capable for tasks involving complex, hierarchically structured information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk 等NeurIPS 2025 · 被引用 27 次
- Hyperbolic Busemann Neural NetworksZiheng Chen, Bernhard Schölkopf, Nicu SebeCVPR 2026 · 被引用 4 次
- Intrinsic Lorentz Neural NetworkXianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu SebeICLR 2026 · 被引用 3 次
- HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented GenerationHiren Madhu, Ngoc Bui, Ali Maatouk, Leandros Tassiulas 等ICML 2026 · 被引用 1 次
- Geometric Constraints for Small Language Models to Understand and Expand Scientific TaxonomiesLiri Fang, Dongqi Fu, Jiawei Han, Jingrui He 等ICLR 2026
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
相关 Paper
- Large Language Models Enhanced Hyperbolic Space Recommender SystemsWentao Cheng, Zhida Qin, Zexue Wu, Pengzhan Zhou 等SIGIR 2025 · 被引用 6 次
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie 等ICLR 2021 · 被引用 19 次
- HyperMiner: Topic Taxonomy Mining with Hyperbolic EmbeddingYishi Xu, Dongsheng Wang, Bo Chen, Ruiying Lu 等NeurIPS 2022 · 被引用 38 次
- Language Models as Hierarchy EncodersYuan He, Moy Yuan, Jiaoyan Chen, Ian HorrocksNeurIPS 2024 · 被引用 37 次
- Complex Hyperbolic Knowledge Graph Embeddings with Fast Fourier TransformHuiru Xiao, Xin Liu, Yangqiu Song, Ginny Y. Wong 等EMNLP 2022 · 被引用 9 次
