SceneForge: Enhancing 3D-text alignment with Structured Scene Compositions
Cristian Sbrolli, Matteo Matteucci
Abstract
The whole is greater than the sum of its parts, even in 3D-text contrastive learning. We introduce SCENEFORGE, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured multi-object scene compositions. SCENEFORGE leverages individual 3D shapes to construct multiobject scenes with explicit spatial relations, pairing them with coherent multi-object descriptions refined by a large language model. By augmenting contrastive training with these structured, compositional samples, SCENEFORGE effectively addresses the scarcity of large-scale 3D-text datasets, significantly enriching data complexity and diversity. We systematically investigate critical design elements, such as the optimal number of objects per scene, the proportion of compositional samples in training batches, and scene construction strategies. Extensive experiments demonstrate that SCENEFORGE delivers substantial performance gains across multiple tasks, including zero-shot classification on ModelNet, ScanObjNN, Objaverse-LVIS, and ScanNet, as well as few-shot part segmentation on ShapeNetPart. SCENEFORGE's compositional augmentations are model-agnostic, consistently improving performance across multiple encoder architectures. Moreover, SCENEFORGE improves 3D visual question answering on ScanQA, generalizes robustly to retrieval scenarios with increasing scene complexity, and showcases spatial reasoning capabilities by adapting spatial configurations to align precisely with textual instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70e09584-e379-4eff-8de3-a76c34d2f66aBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu et al.NeurIPS 2023 · 267 citations
- Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-TrainingYipeng Gao, Zeyu Wang, Wei-Shi Zheng, Cihang Xie et al.CVPR 2024
- CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive LearningLianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin et al.CVPR 2024
- LESS: Label-Efficient and Single-Stage Referring 3D SegmentationXuexun Liu, Xiaoxu Xu, Jinlong Li, Qiudan Zhang et al.NeurIPS 2024 · 7 citations
- DSPNet: Dual-vision Scene Perception for Robust 3D Question AnsweringJingzhou Luo, Yang Liu, Weixing Chen, Zhen Li et al.CVPR 2025
