Coneheads: Hierarchy Aware Attention
Albert Tseng, Tao Yu, Toni J. B. Liu, Christopher De Sa
Abstract
Attention networks such as transformers have achieved state-of-the-art performance in many domains. These networks rely heavily on the dot product attention operator, which computes the similarity between two points by taking their inner product. However, the inner product does not explicitly model the complex structural properties of real world datasets, such as hierarchies between data points. To remedy this, we introduce cone attention, a drop-in replacement for dot product attention based on hyperbolic entailment cones. Cone attention associates two points by the depth of their lowest common ancestor in a hierarchy defined by hyperbolic cones, which intuitively measures the divergence of two points and gives a hierarchy aware similarity score. We test cone attention on a wide variety of models and tasks and show that it improves task-level performance over dot product attention and other baselines, and is able to match dot-product attention with significantly fewer parameters. Our results suggest that cone attention is an effective way to capture hierarchical relationships when calculating attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fac4346-e5c4-4b1f-b1d5-4f89547fd5a8Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Modeling Heterogeneous Hierarchies with Relation-specific Hyperbolic ConesYushi Bai, Zhitao Ying, Hongyu Ren, Jure LeskovecNeurIPS 2021 · 84 citations
- H2KGAT: Hierarchical Hyperbolic Knowledge Graph Attention NetworkShen Wang, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang et al.EMNLP 2020
- Language Models as Hierarchy EncodersYuan He, Moy Yuan, Jiaoyan Chen, Ian HorrocksNeurIPS 2024 · 37 citations
- Hierarchy-Aware Multi-Hop Question Answering over Knowledge GraphsJunnan Dong, Qinggang Zhang, Xiao Huang, Keyu Duan et al.WWW 2023 · 46 citations
- Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale ProblemsSaeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito KoishidaNeurIPS 2025 · 2 citations
