Cross-Task Knowledge Transfer for Semi-supervised Joint 3D Grounding and Captioning
Yang Liu, Daizong Liu, Zongming Guo, Wei Hu
Abstract
3D visual grounding is a fundamental yet important task in multimedia understanding, which aims to locate a specific object in a complicated 3D scene semantically according to a text description. However, this task requires a large number of annotations of labeled text-object pairs for training, so the scarcity of annotated data has been a key obstacle in this task. To this end, this paper makes the first attempt to introduce and address a new semi-supervised setting, where only a few text-object labels are provided during training. Considering most scene data has no annotation, we explore a new solution for unlabeled 3D grounding by additionally training and transferring knowledge from a correlated task, i.e., 3D captioning. Our main insight is that 3D grounding and captioning are complementary and can be iteratively trained with unlabeled data to provide object and text contexts for each other with pseudo-label learning. Specifically, we propose a novel 3D Cross-Task Teacher-Student Framework (3D-CTTSF) for joint 3D grounding and captioning in the semi-supervised setting, where each branch contains parallel grounding and captioning modules. We first pre-train the two modules of the teacher branch with limited labeled data for warm-up. Then, we train the student branch to mimic the ability of the teacher model and iteratively update both branches with the unlabeled data. In particular, we transfer the learned knowledge between the grounding and captioning modules across two branches to generate and refine the pseudo-labels of unlabeled data for providing reliable supervision. To further improve the quality of the pseudo-labels, we design a cross-task pseudo-label generation scheme, filtering low-quality pseudo-labels at the detection, captioning, and grounding levels, respectively. Experimental results on various datasets show competitive performances in both tasks compared to previous fully- and weakly-supervised methods, demonstrating the proposed 3D-CTTSF can serve as an effective solution to overcome the data scarcity issue.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
- Seeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene AttacksDaizong Liu, Wei HuCVPR 2025
- Spatial-Spectral Homogeneous Attacks on Physical-World Large Vision-Language ModelsDaizong Liu, Baoquan Chen, Wei HuAAAI 2026
Related papers
- 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point CloudsDaigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng et al.CVPR 2022 · 80 citations
- UniT3D: A Unified Transformer for 3D Dense Captioning and Visual GroundingDave Zhenyu Chen, Ronghang Hu, Xinlei Chen, Matthias Nießner et al.ICCV 2023 · 82 citations
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
- Towards CLIP-Driven Language-Free 3D Visual Grounding via 2D-3D Relational Enhancement and ConsistencyYuqi Zhang, Han Luo, Yinjie LeiCVPR 2024 · 5 citations
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo et al.AAAI 2025 · 12 citations
