CORN: Contact-based Object Representation for Nonprehensile Manipulation of General Unseen Objects
Yoonyoung Cho, Junhyek Han, Yoontae Cho, Beomjoon Kim
Abstract
Nonprehensile manipulation is essential for manipulating objects that are too thin, large, or otherwise ungraspable in the wild. To sidestep the difficulty of contact modeling in conventional modeling-based approaches, reinforcement learning (RL) has recently emerged as a promising alternative. However, previous RL approaches either lack the ability to generalize over diverse object shapes, or use simple action primitives that limit the diversity of robot motions. Furthermore, using RL over diverse object geometry is challenging due to the high cost of training a policy that takes in high-dimensional sensory inputs. We propose a novel contact-based object representation and pretraining pipeline to tackle this. To enable massively parallel training, we leverage a lightweight patch-based transformer architecture for our encoder that processes point clouds, thus scaling our training across thousands of environments. Compared to learning from scratch, or other shape representation baselines, our representation facilitates both time- and data-efficient learning. We validate the efficacy of our overall system by zero-shot transferring the trained policy to novel real-world objects. Code and videos are available at https://sites.google.com/view/contact-non-prehensile.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ab80d20-c766-4a2b-9dee-532166c46c48Cited by top-tier papers2
- DyWA: Dynamics-Adaptive World Action Model for Generalizable Non-Prehensile ManipulationJiangran Lyu, Ziming Li, Xuesong Shi, Chaoyi Xu et al.ICCV 2025 · 2 citations
- DexMove: Learning Tactile-Guided Non-Prehensile Manipulation with Dexterous HandsPei Lin, Yuzhe Huang, Wanlin Li, Chenxi Xiao et al.ICLR 2026
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
Related papers
- RRL: Resnet as representation for Reinforcement LearningRutav M. Shah, Vikash KumarICML 2021 · 129 citations
- UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic GraspingWenbo Wang, Fangyun Wei, Lei Zhou, Xi Chen et al.CVPR 2025
- MetaMorph: Learning Universal Controllers with TransformersAgrim Gupta, Linxi Fan, Surya Ganguli, Li Fei-FeiICLR 2022 · 130 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- SUGAR : Pre-training 3D Visual Representations for RoboticsShizhe Chen, Ricardo Garcia, Ivan Laptev, Cordelia SchmidCVPR 2024
