Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation
Yongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo, Yifan Wen, Mingtong Dai, weixing chen, Ziliang Chen, Lingbo Liu, Guanbin Li, Liang Lin
Abstract
Recent vision-language-action (VLA) models for multi-task robotic manipulation commonly rely on static viewpoints and shared visual encoders, which limit 3D perception and cause task interference, hindering robustness and generalization. In this work, we propose Task-aware Virtual View Exploration (TVVE), a framework designed to overcome these challenges by integrating virtual view exploration with task-specific representation learning. TVVE employs an efficient exploration policy, accelerated by a novel pseudo-environment, to acquire informative views. Furthermore, we introduce a Task-aware Mixture-of-Experts (TaskMoE) visual encoder to disentangle features across different tasks, boosting both representation fidelity and task generalization. By learning to see the world in a task-aware way, TVVE generates more complete and discriminative visual representations, demonstrating significantly enhanced action prediction across a wide array of manipulation challenges. To further validate the robustness and generalization capability of TVVE under out-of-distribution (OOD) settings, we construct a challenging benchmark, RLBench-OG, covering various visual perturbations and camera pose variations. Extensive experiments on RLBench and RLBench-OG show that our TVVE achieves superior performance over state-of-the-art approaches. In real-robot experiments, TVVE demonstrates exceptional performance and generalizes robustly in multiple OOD settings, including visual disturbances and unseen instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dda5ee4-b6d2-4817-9b68-272c29b4d71dBuilds on9
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An et al.ICLR 2026 · 216 citations
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu et al.ICML 2023 · 158 citations
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive ReasoningFanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You et al.ICLR 2026 · 129 citations
- BAKU: An Efficient Transformer for Multi-Task Policy LearningSiddhant Haldar, Zhuoran Peng, Lerrel PintoNeurIPS 2024 · 120 citations
Related papers
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic ScenesZhiYuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du et al.ICLR 2026 · 23 citations
- Learning View-invariant World Models for Visual Robotic ManipulationJing-Cheng Pang, Nan Tang, Kaiyuan Li, Yuting Tang et al.ICLR 2025
- DyGRO-VLA: Cross-Task Scaling of Vision–Language–Action Models via Dynamic Grouped Residual OptimizationSixu Lin, Yunpeng Qing, Litao Liu, Ming Zhou et al.ICML 2026 · 3 citations
- Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action PolicyTianyi Zhang, Haonan Duan, Haoran Hao, Yu Qiao et al.AAAI 2026 · 5 citations
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic ManipulationCui Miao, Tao Chang, Meihan Wu, Hongbin Xu et al.ICCV 2025 · 7 citations
