Hyperbolic Fine-Tuning for Large Language Models
Menglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong, Jiahong Liu, Irwin King, Rex Ying
Abstract
Large language models (LLMs) have demonstrated remarkable performance across various tasks. However, it remains an open question whether the default Euclidean space is the most suitable choice for LLMs. In this study, we investigate the geometric characteristics of LLMs, focusing specifically on tokens and their embeddings. Our findings reveal that token frequency follows a power-law distribution, where high-frequency tokens (e.g., "the," "that") constitute the minority, while low-frequency tokens (e.g., "apple," "dog") constitute the majority. Furthermore, high-frequency tokens cluster near the origin, whereas low-frequency tokens are positioned farther away in the embedding space. Additionally, token embeddings exhibit hyperbolic characteristics, indicating a latent tree-like structure within the embedding space. Motivated by these observations, we propose HypLoRA, an efficient fine-tuning approach that operates in hyperbolic space to exploit these underlying hierarchical structures better. HypLoRA performs low-rank adaptation directly in hyperbolic space, thereby preserving hyperbolic modeling capabilities throughout the fine-tuning process. Extensive experiments across various base models and reasoning benchmarks, specifically arithmetic and commonsense reasoning tasks, demonstrate that HypLoRA substantially improves LLM performance.
Recent advancements suggest that non-Euclidean geometries, particularly hyperbolic spaces [11,14,15,16,17,18,19,20,21,22,23], offer promising alternatives for modeling hierarchical data. Hyperbolic space, distinguished by its negative curvature, is especially well-suited for representing tree-like hierarchical data due to its exponential volume growth and geometric prior. This geometric property makes hyperbolic space particularly capable for tasks involving complex, hierarchically structured information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94665ea5-9c84-487a-a78b-d01d022056ceCited by top-tier papers15
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk et al.NeurIPS 2025 · 27 citations
- Hyperbolic Busemann Neural NetworksZiheng Chen, Bernhard Schölkopf, Nicu SebeCVPR 2026 · 4 citations
- Intrinsic Lorentz Neural NetworkXianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu SebeICLR 2026 · 3 citations
- HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented GenerationHiren Madhu, Ngoc Bui, Ali Maatouk, Leandros Tassiulas et al.ICML 2026 · 1 citation
- Geometric Constraints for Small Language Models to Understand and Expand Scientific TaxonomiesLiri Fang, Dongqi Fu, Jiawei Han, Jingrui He et al.ICLR 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Large Language Models Enhanced Hyperbolic Space Recommender SystemsWentao Cheng, Zhida Qin, Zexue Wu, Pengzhan Zhou et al.SIGIR 2025 · 6 citations
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie et al.ICLR 2021 · 19 citations
- HyperMiner: Topic Taxonomy Mining with Hyperbolic EmbeddingYishi Xu, Dongsheng Wang, Bo Chen, Ruiying Lu et al.NeurIPS 2022 · 38 citations
- Language Models as Hierarchy EncodersYuan He, Moy Yuan, Jiaoyan Chen, Ian HorrocksNeurIPS 2024 · 37 citations
- Complex Hyperbolic Knowledge Graph Embeddings with Fast Fourier TransformHuiru Xiao, Xin Liu, Yangqiu Song, Ginny Y. Wong et al.EMNLP 2022 · 9 citations
