Language Models as Hierarchy Encoders
Yuan He, Moy Yuan, Jiaoyan Chen, Ian Horrocks
Abstract
Interpreting hierarchical structures latent in language is a key limitation of current language models (LMs). While previous research has implicitly leveraged these hierarchies to enhance LMs, approaches for their explicit encoding are yet to be explored. To address this, we introduce a novel approach to re-train transformer encoder-based LMs as Hierarchy Transformer encoders (HiTs), harnessing the expansive nature of hyperbolic space. Our method situates the output embedding space of pre-trained LMs within a Poincaré ball with a curvature that adapts to the embedding dimension, followed by training on hyperbolic clustering and centripetal losses. These losses are designed to effectively cluster related entities (input as texts) and organise them hierarchically. We evaluate HiTs against pre-trained LMs, standard fine-tuned LMs, and several hyperbolic embedding baselines, focusing on their capabilities in simulating transitive inference, predicting subsumptions, and transferring knowledge across hierarchies. The results demonstrate that HiTs consistently outperform all baselines in these tasks, underscoring the effectiveness and transferability of our re-trained hierarchy encoders.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5d9999c-d14e-4822-9e6f-e7c6980d4eebCited by top-tier papers17
- Agent-OM: Leveraging LLM Agents for Ontology MatchingZhangcheng Qiang, Weiqing Wang, Kerry TaylorVLDB 2025 · 34 citations
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 6 citations
- Hierarchical Retrieval: The Geometry and a Pretrain-Finetune RecipeChong You, Rajesh Jayaram, Ananda Theertha Suresh, Robin Nittka et al.NeurIPS 2025 · 4 citations
- The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsKiho Park, Yo Joong Choe, Yibo Jiang, Victor VeitchICLR 2025 · 3 citations
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengCVPR 2026 · 3 citations
Builds on4
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 791 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Probing BERT in Hyperbolic SpacesBoli Chen, Yao Fu, Guangwei Xu, Pengjun Xie et al.ICLR 2021 · 19 citations
Related papers
- HELM: Hyperbolic Large Language Models via Mixture-of-Curvature ExpertsNeil He, Rishabh Anand, Hiren Madhu, Ali Maatouk et al.NeurIPS 2025 · 27 citations
- Hyperbolic Interaction Model for Hierarchical Multi-Label ClassificationBoli Chen, Xin Huang, Lin Xiao, Zixin Cai et al.AAAI 2020 · 78 citations
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalZiwei Wang, Sameera Ramasinghe, Chenchen Hu, Julien Monteil et al.ICCV 2025 · 4 citations
- HIER: Metric Learning Beyond Class Labels via Hierarchical RegularizationSungyeon Kim, Boseung Jeong, Suha KwakCVPR 2023
- HyperMiner: Topic Taxonomy Mining with Hyperbolic EmbeddingYishi Xu, Dongsheng Wang, Bo Chen, Ruiying Lu et al.NeurIPS 2022 · 38 citations
