SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics
Mengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han, Shanyu Rong, Liming Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
Abstract
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven perception actively with robust, viewpoint-invariant execution accordingly. To this end, we propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Central to our approach is a decoupling of camera and manipulation actions, contrary to shared-action-space, and learning in a bottom-up strategy: we first train semantic camera control on our proposed large-scale dataset, then jointly optimizes both action types via hybrid data. To support this, we introduce ActiveViewPose-200K, comprising * Equal contribution † Corresponding author ‡ Project leader 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We further present ActiveManip-Bench, the first benchmark filling the gap to evaluate active manipulation. Extensive experiments in both simulation and real-world settings show that SaPaVe outperforms recent VLA models such as GR00T N1 and π 0 , achieving up to 31.25% higher success rates in real-world tasks. Our results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Please see the project page at https://lmzpai.github.io/SaPaVe.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 099e9933-35a4-4fd0-bc02-0a81b150b13cBuilds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsHuan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang et al.ICCV 2021 · 419 citations
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An et al.ICLR 2026 · 216 citations
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and ManipulationJiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An et al.NeurIPS 2024 · 154 citations
Related papers
- Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and LanguageQiwei Wu, Rui Zhang, Xin Xiang, Tao Li et al.ICML 2026 · 2 citations
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo et al.CVPR 2026 · 6 citations
- ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic ManipulationZhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue et al.CVPR 2026 · 23 citations
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang et al.ICLR 2026 · 17 citations
- RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action UnderstandingBaoli Sun, Ning Wang, Xinzhu Ma, Anqi Zou et al.ICCV 2025
