Hypformer: Exploring Efficient Transformer Fully in Hyperbolic Space
Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, Rex Ying
摘要
Hyperbolic geometry have shown significant potential in modeling complex structured data, particularly those with underlying treelike and hierarchical structures. Despite the impressive performance of various hyperbolic neural networks across numerous domains, research on adapting the Transformer to hyperbolic space remains limited. Previous attempts have mainly focused on modifying selfattention modules in the Transformer. However, these efforts have fallen short of developing a complete hyperbolic Transformer. This stems primarily from: (i) the absence of well-defined modules in hyperbolic space, including linear transformation layers, Layer-Norm layers, activation functions, dropout operations, etc. (ii) the quadratic time complexity of the existing hyperbolic self-attention module w.r.t the number of input tokens, which hinders its scalability. To address these challenges, we propose, Hypformer, a novel hyperbolic Transformer based on the Lorentz model of hyperbolic geometry. In Hypformer, we introduce two foundational blocks that define the essential modules of the Transformer in hyperbolic space. Furthermore, we develop a linear self-attention mechanism in hyperbolic space, enabling hyperbolic Transformer to process billion-scale graph data and long-sequence inputs for the first time. Our experimental results confirm the effectiveness and efficiency of Hypformer across various datasets, demonstrating its potential as an effective and scalable solution for large-scale data representation and large models. CCS Concepts • Computing methodologies → Machine learning; Knowledge representation and reasoning; • Mathematics of computing → Geometric topology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Hyperbolic Fine-Tuning for Large Language ModelsMenglin Yang, Ram Samarth B. B., Aosong Feng, Bo Xiong 等NeurIPS 2025 · 被引用 31 次
- RiemannGFM: Learning a Graph Foundation Model from Riemannian GeometryLi Sun, Zhenhao Huang, Suyang Zhou, Qiqi Wan 等WWW 2025 · 被引用 31 次
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk 等NeurIPS 2025 · 被引用 27 次
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 被引用 15 次
- Hyperbolic Busemann Neural NetworksZiheng Chen, Bernhard Schölkopf, Nicu SebeCVPR 2026 · 被引用 4 次
它引用的顶会 Paper39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- Fully Hyperbolic Convolutional Neural Networks for Computer VisionAhmad Bdeir, Kristian Schwethelm, Niels LandwehrICLR 2024 · 被引用 45 次
- Lorentzian Residual Neural NetworksNeil He, Menglin Yang, Rex YingKDD 2025 · 被引用 1 次
- Fully Hyperbolic Neural NetworksWeize Chen, Xu Han, Yankai Lin, Hexu Zhao 等ACL 2022
- Intrinsic Lorentz Neural NetworkXianglong Shi, Ziheng Chen, Yunhan Jiang, Nicu SebeICLR 2026 · 被引用 3 次
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian 等ICCV 2025 · 被引用 5 次
