Hypformer: Exploring Efficient Transformer Fully in Hyperbolic Space
Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, Rex Ying
Abstract
Hyperbolic geometry have shown significant potential in modeling complex structured data, particularly those with underlying treelike and hierarchical structures. Despite the impressive performance of various hyperbolic neural networks across numerous domains, research on adapting the Transformer to hyperbolic space remains limited. Previous attempts have mainly focused on modifying selfattention modules in the Transformer. However, these efforts have fallen short of developing a complete hyperbolic Transformer. This stems primarily from: (i) the absence of well-defined modules in hyperbolic space, including linear transformation layers, Layer-Norm layers, activation functions, dropout operations, etc. (ii) the quadratic time complexity of the existing hyperbolic self-attention module w.r.t the number of input tokens, which hinders its scalability. To address these challenges, we propose, Hypformer, a novel hyperbolic Transformer based on the Lorentz model of hyperbolic geometry. In Hypformer, we introduce two foundational blocks that define the essential modules of the Transformer in hyperbolic space. Furthermore, we develop a linear self-attention mechanism in hyperbolic space, enabling hyperbolic Transformer to process billion-scale graph data and long-sequence inputs for the first time. Our experimental results confirm the effectiveness and efficiency of Hypformer across various datasets, demonstrating its potential as an effective and scalable solution for large-scale data representation and large models. CCS Concepts • Computing methodologies → Machine learning; Knowledge representation and reasoning; • Mathematics of computing → Geometric topology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa4e5252-5d9d-4c88-85ad-7991290f3242Cited by top-tier papers14
- Hyperbolic Fine-Tuning for Large Language ModelsMenglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong et al.NeurIPS 2025 · 31 citations
- RiemannGFM: Learning a Graph Foundation Model from Riemannian GeometryLi Sun, Zhenhao Huang, Suyang Zhou, Qiqi Wan et al.WWW 2025 · 31 citations
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk et al.NeurIPS 2025 · 27 citations
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 15 citations
- Hyperbolic Busemann Neural NetworksZiheng Chen, Bernhard Schölkopf, Nicu SebeCVPR 2026 · 4 citations
Builds on39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Fully Hyperbolic Convolutional Neural Networks for Computer VisionAhmad Bdeir, Kristian Schwethelm, Niels LandwehrICLR 2024 · 45 citations
- Lorentzian Residual Neural NetworksNeil He, Menglin Yang, Rex YingKDD 2025 · 1 citation
- Fully Hyperbolic Neural NetworksWeize Chen, Xu Han, Yankai Lin, Hexu Zhao et al.ACL 2022
- Intrinsic Lorentz Neural NetworkXianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu SebeICLR 2026 · 3 citations
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian et al.ICCV 2025 · 5 citations
