SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics
Mengzhen Liu, Enshen Zhou, Cheng Chi, Yi Han, Shanyu Rong, Liming Chen, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang
摘要
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven perception actively with robust, viewpoint-invariant execution accordingly. To this end, we propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Central to our approach is a decoupling of camera and manipulation actions, contrary to shared-action-space, and learning in a bottom-up strategy: we first train semantic camera control on our proposed large-scale dataset, then jointly optimizes both action types via hybrid data. To support this, we introduce ActiveViewPose-200K, comprising * Equal contribution † Corresponding author ‡ Project leader 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We further present ActiveManip-Bench, the first benchmark filling the gap to evaluate active manipulation. Extensive experiments in both simulation and real-world settings show that SaPaVe outperforms recent VLA models such as GR00T N1 and π 0 , achieving up to 31.25% higher success rates in real-world tasks. Our results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Please see the project page at https://lmzpai.github.io/SaPaVe.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsHuan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang 等ICCV 2021 · 被引用 419 次
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An 等ICLR 2026 · 被引用 216 次
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and ManipulationJiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An 等NeurIPS 2024 · 被引用 154 次
相关 Paper
- Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and LanguageQiwei Wu, Rui Zhang, Xin Xiang, Tao Li 等ICML 2026 · 被引用 2 次
- Learning to See and Act: Task-Aware Virtual View Exploration for Robotic ManipulationYongjie Bai, Zhouxia Wang, Yang Liu, Kaijun Luo 等CVPR 2026 · 被引用 6 次
- ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic ManipulationZhenyang Liu, Yongchong Gu, Yikai Wang, Xiangyang Xue 等CVPR 2026 · 被引用 23 次
- RoboInter: A Holistic Intermediate Representation Suite Towards Robotic ManipulationHao Li, Ziqin Wang, Zi-han Ding, Shuai Yang 等ICLR 2026 · 被引用 17 次
- RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action UnderstandingBaoli Sun, Ning Wang, Xinzhu Ma, Anqi Zou 等ICCV 2025
