Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
Wenrui Li, Wei Han, Yandu Chen, Yeyu Chai, Yidan Lu, Xingtao Wang, Xiaopeng Fan
Abstract
Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemannbased Multi-scale Attention Reasoning Network (RMARN) for text-3D retrieval. Specifically, the extracted text and point cloud features are refined by their respective Adaptive Feature Refiner (AFR). Furthermore, we introduce the innovative Riemann Local Similarity (RLS) module and the Global Pooling Similarity (GPS) module. However, as 3D point cloud data and text data often possess complex geometric structures in high-dimensional space, the proposed RLS employs a novel Riemann Attention Mechanism to reflect the intrinsic geometric relationships of the data. Without explicitly defining the manifold, RMARN learns the manifold parameters to better represent the distances between text-point cloud samples. To address the challenges of lacking paired text-3D data, we have created the large-scale Text-3D Retrieval dataset T3DR-HIT, which comprises over 3,380 pairs of text and point cloud data. T3DR-HIT contains coarse-grained indoor 3D scenes and fine-grained Chinese artifact scenes, consisting of 1,380 and over 2,000 text-3D pairs, respectively. Experiments on our custom datasets demonstrate the superior performance of the proposed method. Our code and proposed datasets are available at https://github.com/liwrui/RMARN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdbcaafc-4ba7-4b7e-990e-0d8d1f803d2cCited by top-tier papers1
Ask how each one uses itBuilds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Language-Agnostic Visual-Semantic EmbeddingsJonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, Rodrigo C. BarrosICCV 2019 · 56 citations
- Offline and Online Optical Flow Enhancement for Deep Video CompressionChuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang et al.AAAI 2024 · 35 citations
Related papers
- Hyperbolic Hierarchical Alignment Reasoning Network for Text-3D RetrievalWenrui Li, Yidan Lu, Yeyu Chai, Rui Zhao et al.AAAI 2026
- Context-aware Alignment and Mutual Masking for 3D-Language Pre-trainingZhao Jin, Munawar Hayat, Yuwei Yang, Yulan Guo et al.CVPR 2023
- TVDRNet: Text-driven Viewpoint Optimization via Differentiable Rendering for 3D Reasoning SegmentationTingran Wang, Changshuo Wang, Pinjie Xu, ZhangHuang et al.ICML 2026
- 3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at ScaleYijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang et al.AAAI 2026 · 12 citations
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.EMNLP 2023 · 7 citations
