SUGAR : Pre-training 3D Visual Representations for Robotics
Shizhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia Schmid
摘要
Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet, prevailing approaches focus on pre-training 2D representations, being sub-optimal to deal with occlusions and accurately localize objects in complex 3D scenes. Meanwhile, 3D representation learning has been limited to single-object understanding. To address these limitations, we introduce a novel 3D pre-training framework for robotics named SUGAR that captures semantic, geometric and affordance properties of objects through 3D point clouds. We underscore the importance of cluttered scenes in 3D representation learning, and automatically construct a multi-object dataset benefiting from cost-free supervision in simulation. SUGAR employs a versatile transformer-based model to jointly address five pre-training tasks, namely cross-modal knowledge distillation for semantic learning, masked point modeling to understand geometry structures, grasping pose synthesis for object affordance, 3D instance segmentation and referring expression grounding to analyze cluttered scenes. We evaluate our learned representation on three robotic-related tasks, namely, zero-shot 3D object recognition, referring expression grounding, and language-driven robotic manipulation. Experimental results show that SUGAR's 3D representation outperforms state-of-the-art 2D and 3D representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou 等NeurIPS 2024 · 被引用 25 次
- ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian SplattingJiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long 等CVPR 2026 · 被引用 7 次
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang 等CVPR 2026 · 被引用 6 次
- 4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware BootstrappingXindan Zhang, Weilong Yan, YUFEI SHI, Xuerui Qiu 等ICML 2026 · 被引用 6 次
- Dynamic Focused Masking for Autoregressive Embodied Occupancy PredictionYuan Sun, Julio Contreras, Jorge OrtizNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai 等NeurIPS 2022 · 被引用 1,270 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
相关 Paper
- GEAL: Generalizable 3D Affordance Learning with Cross-Modal ConsistencyDongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee LeeCVPR 2025
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- 4D Visual Pre-Training for Robot LearningChengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao 等ICCV 2025 · 被引用 1 次
- UniGS: Unified Language-Image-3D Pretraining with Gaussian SplattingHaoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song 等ICLR 2025
- CORN: Contact-based Object Representation for Nonprehensile Manipulation of General Unseen ObjectsYoonyoung Cho, Junhyek Han, Yoontae Cho, Beomjoon KimICLR 2024 · 被引用 20 次
