SceneForge: Enhancing 3D-text alignment with Structured Scene Compositions
Cristian Sbrolli, Matteo Matteucci
摘要
The whole is greater than the sum of its parts, even in 3D-text contrastive learning. We introduce SCENEFORGE, a novel framework that enhances contrastive alignment between 3D point clouds and text through structured multi-object scene compositions. SCENEFORGE leverages individual 3D shapes to construct multiobject scenes with explicit spatial relations, pairing them with coherent multi-object descriptions refined by a large language model. By augmenting contrastive training with these structured, compositional samples, SCENEFORGE effectively addresses the scarcity of large-scale 3D-text datasets, significantly enriching data complexity and diversity. We systematically investigate critical design elements, such as the optimal number of objects per scene, the proportion of compositional samples in training batches, and scene construction strategies. Extensive experiments demonstrate that SCENEFORGE delivers substantial performance gains across multiple tasks, including zero-shot classification on ModelNet, ScanObjNN, Objaverse-LVIS, and ScanNet, as well as few-shot part segmentation on ShapeNetPart. SCENEFORGE's compositional augmentations are model-agnostic, consistently improving performance across multiple encoder architectures. Moreover, SCENEFORGE improves 3D visual question answering on ScanQA, generalizes robustly to retrieval scenarios with increasing scene complexity, and showcases spatial reasoning capabilities by adapting spatial configurations to align precisely with textual instructions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu 等NeurIPS 2023 · 被引用 267 次
- Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-TrainingYipeng Gao, Zeyu Wang, Wei-Shi Zheng, Cihang Xie 等CVPR 2024
- CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive LearningLianggangxu Chen, Xuejiao Wang, Jiale Lu, Shaohui Lin 等CVPR 2024
- LESS: Label-Efficient and Single-Stage Referring 3D SegmentationXuexun Liu, Xiaoxu Xu, Jinlong Li, Qiudan Zhang 等NeurIPS 2024 · 被引用 7 次
- DSPNet: Dual-vision Scene Perception for Robust 3D Question AnsweringJingzhou Luo, Yang Liu, Weixing Chen, Zhen Li 等CVPR 2025
