Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
Minghuan Liu, Zhengbang Zhu, Xiaoshen Han, Peng Hu, Haotong Lin, Xinyao Li, Jingxiao Chen, Jiafeng Xu, Yichu Yang, Yunfeng Lin, Xinghang Li, Yong Yu
Abstract
Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance, size, and shape-than on texture when interacting with objects. Since such 3D geometric information can be acquired from widely available depth cameras, it appears feasible to endow robots with similar perceptual capabilities. Our pilot study found that using depth cameras for manipulation is challenging, primarily due to their limited accuracy and susceptibility to various types of noise. In this work, we propose Camera Depth Models (CDMs) as a simple plugin on daily-use depth cameras, which take RGB images and raw depth signals as input and output denoised, accurate metric depth. To achieve this, we develop a neural data engine that generates high-quality paired data from simulation by modeling a depth camera's noise pattern. Our results show that CDMs achieve nearly simulation-level accuracy in depth prediction, effectively bridging the sim-to-real gap for manipulation tasks. Notably, our experiments demonstrate, for the first time, that a policy trained on raw simulated depth, without the need for adding noise or real-world fine-tuning, generalizes seamlessly to real-world robots on two challenging long-horizon tasks involving articulated, reflective, and slender objects, with little to no performance degradation. We hope our findings will inspire future research in utilizing simulation data and 3D information in general robot policies. We release the dataset, models for various depth cameras, along with an easy-to-use guide for sim-to-real transfer at https://manipulation-as-in-simulation.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6def592f-9bb2-431a-bd70-6e5e5fca671fCited by top-tier papers4
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation PriorsZhengshen Zhang, Hao Li, Yalun Dai, Zhengbang Zhu et al.ICLR 2026 · 27 citations
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
- Zero-Shot Depth Completion with Vision-Language ModelZhiqiang Yan, Yuan Wu, Gim Hee LeeCVPR 2026 · 1 citation
- PRISM: Learning Realistic Depth via Physics-Grounded Noise Disentanglement with Semantic-Geometric CollaborationShawn Leung, Jiacheng Liu, Mingyang Sun, Qichen He et al.ICML 2026
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual ManipulationChongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan et al.CVPR 2026 · 10 citations
- Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control InterfaceYujie Zhao, Hongwei Fan, Di Chen, Shengcong Chen et al.CVPR 2026 · 8 citations
- SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Yang Yue et al.AAAI 2026 · 2 citations
- DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics ModelXueyi Liu, He Wang, Li YiICLR 2026 · 26 citations
- Physics-based Differentiable Depth Sensor SimulationBenjamin Planche, Rajat Vikram SinghICCV 2021 · 11 citations
