SUGAR : Pre-training 3D Visual Representations for Robotics
Shizhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia Schmid
Abstract
Learning generalizable visual representations from Internet data has yielded promising results for robotics. Yet, prevailing approaches focus on pre-training 2D representations, being sub-optimal to deal with occlusions and accurately localize objects in complex 3D scenes. Meanwhile, 3D representation learning has been limited to single-object understanding. To address these limitations, we introduce a novel 3D pre-training framework for robotics named SUGAR that captures semantic, geometric and affordance properties of objects through 3D point clouds. We underscore the importance of cluttered scenes in 3D representation learning, and automatically construct a multi-object dataset benefiting from cost-free supervision in simulation. SUGAR employs a versatile transformer-based model to jointly address five pre-training tasks, namely cross-modal knowledge distillation for semantic learning, masked point modeling to understand geometry structures, grasping pose synthesis for object affordance, 3D instance segmentation and referring expression grounding to analyze cluttered scenes. We evaluate our learned representation on three robotic-related tasks, namely, zero-shot 3D object recognition, referring expression grounding, and language-driven robotic manipulation. Experimental results show that SUGAR's 3D representation outperforms state-of-the-art 2D and 3D representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8808a8ed-a68b-4f31-82d9-f32f2e4bb792Cited by top-tier papers13
- A Unified Framework for 3D Scene UnderstandingWei Xu, Chunsheng Shi, Sifan Tu, Xin Zhou et al.NeurIPS 2024 · 25 citations
- ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian SplattingJiayu Ding, Xinpeng Liu, Zhiyi Pan, Shiqiang Long et al.CVPR 2026 · 7 citations
- Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot LearningQiwei Liang, Boyang Cai, Minghao Lai, Sitong Zhuang et al.CVPR 2026 · 6 citations
- 4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware BootstrappingXindan Zhang, Weilong Yan, YUFEI SHI, Xuerui Qiu et al.ICML 2026 · 6 citations
- Dynamic Focused Masking for Autoregressive Embodied Occupancy PredictionYuan Sun, Julio Contreras, Jorge OrtizNeurIPS 2025 · 3 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
Related papers
- GEAL: Generalizable 3D Affordance Learning with Cross-Modal ConsistencyDongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee LeeCVPR 2025
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- 4D Visual Pre-Training for Robot LearningChengkai Hou, Yanjie Ze, Yankai Fu, Zeyu Gao et al.ICCV 2025 · 1 citation
- UniGS: Unified Language-Image-3D Pretraining with Gaussian SplattingHaoyuan Li, Yanpeng Zhou, Tao Tang, Jifei Song et al.ICLR 2025
- CORN: Contact-based Object Representation for Nonprehensile Manipulation of General Unseen ObjectsYoonyoung Cho, Junhyek Han, Yoontae Cho, Beomjoon KimICLR 2024 · 20 citations
