Multi-view Invariance Learning for 3D Scene Graph Pre-training via Collaborative Cross-Modal Regularization
Yucheng Huang, Luping Ji, Ruijie Xiao, Jiayuan Sun
Abstract
3D scene graph generation is a pivotal task in scene understanding. Its performance is easy to be constrained by the limited availability of annotated data. Currently, the existing solutions on point cloud pre-training usually emphasize on object-centric representations while neglecting the predicate feature learning. This limitation significantly hinders their relational reasoning capabilities, as inter-object relationships are fundamentally governed by predicate features. To enhance 3D Scene Graphs Pre-training, this paper proposes a task-specific Multi-view Invariance Learning framework with Collaborative Cross-modal Regularization. In detail, the inherent horizontal-rotation invariance of 3D objects and their semantic relationships are leveraged to construct a self-supervised paradigm for triplet feature learning. Moreover, our framework harnesses the cross-modal prior knowledge from the vision-language model to regularize model optimization. It could further achieve the semantic discrimination via unsupervised deep clustering. To resolve the knowledge discrepancies arising from the pre-trained model in fine-tuning, a predicate adapter equipped with knowledge filtering gate is devised to selectively aggregate the predicate features of pre-trained model. Extensive experiments demonstrate that our framework is effective in boosting 3D scene graph generation performance, surpassing state-of-the-art ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 333 citations
- Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingZekun Qi, Runpei Dong, Guofan Fan, Zheng Ge et al.ICML 2023 · 209 citations
Related papers
- CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive LearningLianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin et al.CVPR 2024
- Knowledge-inspired 3D Scene Graph Prediction in Point CloudShoulong Zhang, Shuai Li, Aimin Hao, Hong QinNeurIPS 2021 · 54 citations
- VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point CloudZiqin Wang, Bowen Cheng, Lichen Zhao, Dong Xu et al.CVPR 2023
- Vision-Language Pre-training with Object Contrastive Learning for 3D Scene UnderstandingTaolin Zhang, Sunan He, Tao Dai, Zhi Wang et al.AAAI 2024 · 42 citations
- Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant FeaturesWenhuan Huang, Yi Ji, Guiqian Zhu, Li Ying et al.CVPR 2025
