Learning View-invariant World Models for Visual Robotic Manipulation
Jing-Cheng Pang, Nan Tang, Kaiyuan Li, Yuting Tang, Xin-Qiang Cai, Zhen-Yu Zhang, Gang Niu, Masashi Sugiyama, Yang Yu
Abstract
Robotic manipulation tasks often rely on visual inputs from cameras to perceive the environment. However, previous approaches still suffer from performance degradation when the camera's viewpoint changes during manipulation. In this paper, we propose ReViWo (Representation learning for View-invariant World model), leveraging multi-view data to learn robust representations for control under viewpoint disturbance. ReViWo utilizes an autoencoder framework to reconstruct target images by an architecture that combines view-invariant representation (VIR) and view-dependent representation. To train ReViWo, we collect multi-view data in simulators with known view labels. Meanwhile, ReViWo is simutaneously trained on Open X-Embodiment datasets without view labels. The VIR is then used to train a world model on pre-collected manipulation data and a policy through interaction with the world model. We evaluate the effectiveness of ReViWo in various viewpoint disturbance scenarios, including control under novel camera positions and frequent camera shaking, using the Meta-world & PandaGym environments. Besides, we also conduct experiments on real world ALOHA robot. The results demonstrate that ReViWo maintains robust performance under viewpoint disturbance, while baseline methods suffer from significant performance degradation. Furthermore, we show that the VIR captures taskrelevant state information and remains stable for observations from novel viewpoints, validating the efficacy of the ReViWo approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial ModelingWeiqi Li, Quande Zhang, Ruifeng Zhai, Liang Lin et al.CVPR 2026 · 12 citations
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint ReconstructionJung Min Lee, Dohyeok Lee, Seokhun Ju, Taehyun Cho et al.ICML 2026 · 9 citations
- Learning to Act Robustly with View-Invariant Latent ActionsYoungjoon Jeong, Junha Chun, Taesup KimCVPR 2026 · 5 citations
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng et al.ICLR 2026 · 5 citations
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic ManipulationNan Tang, Jing-Cheng Pang, Guanlin Li, Chao Qian et al.ICML 2026 · 1 citation
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- Multi-View Masked World Models for Visual Robotic ManipulationYounggyo Seo, Junsu Kim, Stephen James, Kimin Lee et al.ICML 2023 · 99 citations
- DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic ManipulationKaizhao Zhang, Tian Niu, Tianyu Liu, Chenen Guo et al.CVPR 2026
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo et al.CVPR 2026 · 6 citations
- Vision-Based Manipulators Need to Also See from Their HandsKyle Hsu, Moo Jin Kim, Rafael Rafailov, Jiajun Wu et al.ICLR 2022 · 60 citations
- DRIBO: Robust Deep Reinforcement Learning via Multi-View Information BottleneckJiameng Fan, Wenchao LiICML 2022 · 49 citations
