X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioning
Zhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo, Guanbin Li, Shuguang Cui, Zhen Li
Abstract
3D dense captioning aims to describe individual objects in 3D scenes by natural language, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though aggregating 2D features into point clouds may be beneficial, it introduces an extra computational burden, especially in the inference phase. In this study, we investigate a cross-modal knowledge transfer using Transformer for 3D dense captioning, namely X-Trans2Cap. Our proposed X-Trans2Cap effectively boost the performance of single-modal 3D captioning through the knowledge distillation enabled by a teacher-student framework. In practice, during the training phase, the teacher network exploits auxiliary 2D modality and guides the student network that only takes point clouds as input through the feature consistency constraints. Owing to the well-designed cross-modal feature fusion module and the feature alignment in the training phase, X-Trans2Cap acquires rich appearance information embedded in 2D images with ease. Thus, a more faithful caption can be generated only using point clouds during the inference. Qualitative and quantitative results confirm that X-Trans2Cap outperforms previous state-of-the-art by a large margin, i.e., about +21 and +16 CIDEr points on ScanRefer and Nr3D datasets, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 735c7243-32a4-46a0-9ab9-879010df3dc6Cited by top-tier papers35
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersHaifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang et al.NeurIPS 2024 · 230 citations
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang et al.ICLR 2026 · 121 citations
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner et al.ICCV 2023 · 82 citations
- Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisXu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao et al.NeurIPS 2022 · 49 citations
- Graph Enhanced Contrastive Learning for Radiology Findings SummarizationJinpeng Hu, Zhuo Li, Zhihong Chen, Zhen Li et al.ACL 2022 · 40 citations
Builds on14
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- SAT: 2D Semantics Assisted Training for 3D Visual GroundingZhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo LuoICCV 2021 · 166 citations
- Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsChaoda Zheng, Xu Yan, Jiantao Gao, Weibing Zhao et al.ICCV 2021 · 116 citations
Related papers
- Scan2Cap: Context-Aware Dense Captioning in RGB-D ScansDave Zhenyu Chen, Ali Gholami, Matthias Nießner, Angel X. ChangCVPR 2021
- End-to-End 3D Dense Captioning with Vote2Cap-DETRSijin Chen, Hongyuan Zhu, Xin Chen, Yinjie Lei et al.CVPR 2023
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual GroundingDailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui et al.ACM MM 2021 · 81 citations
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.NeurIPS 2022 · 173 citations
