Language-Assisted 3D Feature Learning for Semantic Scene Understanding
Junbo Zhang, Guofan Fan, Guanghan Wang, Zhengyuan Su, Kaisheng Ma, Li Yi
摘要
Learning descriptive 3D features is crucial for understanding 3D scenes with diverse objects and complex structures. However, it is usually unknown whether important geometric attributes and scene context obtain enough emphasis in an end-to-end trained 3D scene understanding network. To guide 3D feature learning toward important geometric attributes and scene context, we explore the help of textual scene descriptions. Given some free-form descriptions paired with 3D scenes, we extract the knowledge regarding the object relationships and object attributes. We then inject the knowledge to 3D feature learning through three classification-based auxiliary tasks. This language-assisted training can be combined with modern object detection and instance segmentation methods to promote 3D semantic scene understanding, especially in a label-deficient regime. Moreover, the 3D feature learned with language assistance is better aligned with the language features, which can benefit various 3D-language multimodal tasks. Experiments on several benchmarks of 3D-only and 3D-language tasks demonstrate the effectiveness of our language-assisted 3D feature learning. Code is available at https://github.com/Asterisci/Language-Assisted-3D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi 等ICLR 2024 · 被引用 315 次
- VPP: Efficient Conditional 3D Generation via Voxel-Point Progressive RepresentationZekun Qi, Muzhou Yu, Runpei Dong, Kaisheng MaNeurIPS 2023 · 被引用 22 次
- Autoencoders as Cross-Modal Teachers: Can Pretrained 2D Image Transformers Help 3D Representation Learning?Runpei Dong, Zekun Qi, Linfeng Zhang, Junbo Zhang 等ICLR 2023 · 被引用 21 次
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 被引用 12 次
- IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma 等AAAI 2025 · 被引用 5 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak 等ICCV 2019 · 被引用 474 次
相关 Paper
- All in One: Visual-Description-Guided Unified Point Cloud SegmentationZongyan Han, Mohamed El Amine Boudjoghra, Jiahua Dong, Jinhong Wang 等ICCV 2025 · 被引用 1 次
- Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal VisionXiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah SnavelyICCV 2021 · 被引用 26 次
- LLaVA-3D: A Simple Yet Effective Pathway to Empowering LMMs with 3D CapabilitiesChenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang 等ICCV 2025 · 被引用 24 次
- 3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point CloudMingtao Feng, Haoran Hou, Liang Zhang, Zijie Wu 等CVPR 2023
- VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point CloudZiqin Wang, Bowen Cheng, Lichen Zhao, Dong Xu 等CVPR 2023
