Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
Ziyao He, Yingjie Liu, Yangrui Zhang, Mingsong Chen, Xuan Tang, Xian Wei
Abstract
Accurate 3D scene description is fundamental to robotic navigation and augmented reality, yet current dense captioning methods face significant limitations in processing sparse point cloud data. % Existing approaches that apply Euclidean embedding spaces struggle to simultaneously preserve fine-grained local geometric details and model exponentially growing global semantic hierarchies, leading to either inaccurate localization or disjointed, shallow scene descriptions. % In this work, we propose a novel Curvature-Aware Captioning framework, integrating novel non-Euclidean geodesic attention mechanisms, to resolve the localization-contextualization conflict. % Specifically, self-attention within Oblique space enforces dimensional homogeneity while establishing long-range dependencies. Bidirectional geodesic cross-attention within Lorentz space models hierarchical semantic relationships across scene instances, enabling simultaneous precision in object localization and coherence in scene descriptions. % Theoretical analysis confirms that the curvature complementarity between the Oblique manifold and Lorentz hyperboloid resolves the Euclidean-hyperbolic conflict, ensuring feature stability via isotropic optimization while preserving inherent hierarchical relationships. Extensive experiments on ScanRefer and Nr3D benchmarks demonstrate state-of-the-art performance, with significant gains in both localization accuracy and descriptive richness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb9d09ee-58f0-497b-86cb-72936d409fc9Builds on23
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Hyperbolic Neural Networks++Ryohei Shimizu, Yusuke Mukuta, Tatsuya HaradaICLR 2021 · 791 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
Related papers
- Scan2Cap: Context-Aware Dense Captioning in RGB-D ScansDave Zhenyu Chen, Ali Gholami, Matthias Nießner, Angel X. ChangCVPR 2021
- PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic RectificationXinxing Yu, Liying Yang, Hao Mo, Hui Ma et al.ICML 2026
- Geodesic Self-Attention for 3D Point CloudsZhengyu Li, Xuan Tang, Zihao Xu, Xihao Wang et al.NeurIPS 2022 · 18 citations
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.EMNLP 2023 · 7 citations
- EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual GroundingGwangWook Park, Hyo-Jun Lee, Jong-Hyeon Baek, Hanul Kim et al.CVPR 2026
