Vision-Based Manipulators Need to Also See from Their Hands
Kyle Hsu, Moo Jin Kim, Rafael Rafailov, Jiajun Wu, Chelsea Finn
Abstract
We study how the choice of visual perspective affects learning and generalization in the context of physical manipulation from raw sensor observations. Compared with the more commonly used global third-person perspective, a hand-centric (eye-in-hand) perspective affords reduced observability, but we find that it consistently improves training efficiency and out-of-distribution generalization. These benefits hold across a variety of learning algorithms, experimental settings, and distribution shifts, and for both simulated and real robot apparatuses. However, this is only the case when hand-centric observability is sufficient; otherwise, including a third-person perspective is necessary for learning, but also harms outof-distribution generalization. To mitigate this, we propose to regularize the thirdperson information stream via a variational information bottleneck. On six representative manipulation tasks with varying hand-centric observability adapted from the Meta-World benchmark, this results in a state-of-the-art reinforcement learning agent operating from both perspectives improving its out-of-distribution generalization on every task. While some practitioners have long put cameras in the hands of robots, our work systematically analyzes the benefits of doing so and provides simple and broadly applicable insights for improving end-to-end learned vision-based robotic manipulation. 1 Figure 1: Illustration suggesting the role that visual perspective can play in facilitating the acquisition of symmetries with respect to certain transformations on the world state s. T0: planar translation of the end-effector and cube. T1: vertical translation of the table surface, end-effector, and cube. T2: addition of distractor objects. O3: third-person perspective. O h : hand-centric perspective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c97c9c3-3d2c-46df-a5e6-7bec92b3a904Cited by top-tier papers8
- Multi-View Masked World Models for Visual Robotic ManipulationYounggyo Seo, Junsu Kim, Stephen James, Kimin Lee et al.ICML 2023 · 99 citations
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 18 citations
- The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement LearningMoritz Schneider, Robert Krug, Narunas Vaskevicius, Luigi Palmieri et al.NeurIPS 2024 · 10 citations
- A Practical Guide for Incorporating Symmetry in Diffusion PolicyDian Wang, Boce Hu, Shuran Song, Robin Walters et al.NeurIPS 2025 · 9 citations
- Real-World Reinforcement Learning of Active Perception BehaviorsEdward S. Hu, Jie Wang, Xingfang Yuan, Fiona Luo et al.NeurIPS 2025 · 6 citations
Builds on5
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto et al.NeurIPS 2020 · 833 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement LearningDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICLR 2022 · 457 citations
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark et al.ICLR 2020 · 109 citations
Related papers
- Learning View-invariant World Models for Visual Robotic ManipulationJing-Cheng Pang, Nan Tang, Kaiyuan Li, Yuting Tang et al.ICLR 2025
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action PolicyTianyi Zhang, Haonan Duan, Haoran Hao, Yu Qiao et al.AAAI 2026 · 5 citations
- 3D-aware Disentangled Representation for Compositional Reinforcement LearningSungbin Mun, Younghwan Lee, Cheolhui MIn, Mineui Hong et al.ICLR 2026
- EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual AlignmentWei Feng, Chi Zhang, Nan Li, Qian Zhang et al.CVPR 2026
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo et al.CVPR 2026 · 6 citations
