Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D Retrieval
Wenrui Li, Yidan Lu, Yeyu Chai, Rui Zhao, Hengyu Man, Xiaopeng Fan
Abstract
With the daily influx of 3D data on the internet, text-3D retrieval has gained increasing attention. However, current methods face two major challenges: Hierarchy Representation Collapse (HRC) and Redundancy-Induced Saliency Dilution (RISD). HRC compresses abstract-to-specific and whole-topart hierarchies in Euclidean embeddings, while RISD averages noisy fragments, obscuring critical semantic cues and diminishing the model's ability to distinguish hard negatives. To address these challenges, we introduce the Hyperbolic Hierarchical Alignment Reasoning Network (H 2 ARN) for text-3D retrieval. H 2 ARN embeds both text and 3D data in a Lorentz-model hyperbolic space, where exponential volume growth inherently preserves hierarchical distances. A hierarchical ordering loss constructs a shrinking entailment cone around each text vector, ensuring that the matched 3D instance falls within the cone, while an instance-level contrastive loss jointly enforces separation from non-matching samples. To tackle RISD, we propose a contribution-aware hyperbolic aggregation module that leverages Lorentzian distance to assess the relevance of each local feature and applies contribution-weighted aggregation guided by hyperbolic geometry, enhancing discriminative regions while suppressing redundancy without additional supervision. We also release the expanded T3DR-HIT v2 benchmark, which contains 8,935 text-to-3D pairs, 2.6 times the original size, covering both fine-grained cultural artefacts and complex indoor scenes. Our codes are available at https://github.com/liwrui/H2ARN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c1562d8-8e6f-4e8d-9c51-ccc45fbaa5e8Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Hierarchical Scene Normality-Binding Modeling for Anomaly Detection in Surveillance VideosQianyue Bao, Fang Liu, Yang Liu, Licheng Jiao et al.ACM MM 2022 · 50 citations
- The Devil is in the Crack Orientation: A New Perspective for Crack DetectionZhuangzhuang Chen, Jin Zhang, Zhuonan Lai, Guanming Zhu et al.ICCV 2023 · 26 citations
Related papers
- Riemann-based Multi-scale Attention Reasoning Network for Text-3D RetrievalWenrui Li, Wei Han, Yandu Chen, Yeyu Chai et al.AAAI 2025 · 6 citations
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian et al.ICCV 2025 · 5 citations
- HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented GenerationHiren Madhu, Ngoc Bui, Ali Maatouk, Leandros Tassiulas et al.ICML 2026 · 1 citation
- HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape GenerationZhiying Leng, Tolga Birdal, Xiaohui Liang, Federico TombariCVPR 2024
- Hyperbolic Gramian Volumes for Multimodal AlignmentSaiyang Na, Feng Jiang, Qifeng Zhou, Wenliang Zhong et al.CVPR 2026
