SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
Qianzhong Chen, Justin Yu, Mac Schwager, Pieter Abbeel, Fred Shentu, Philipp Wu
摘要
Large-scale robot learning has made progress on complex manipulation tasks, yet long-horizon, contact-rich problems—especially those involving deformable objects—remain challenging due to inconsistent demonstration quality. We propose a stage-aware, video-based reward modeling framework that jointly predicts task stage and fine-grained progress, using natural-language subtask annotations to derive consistent labels across variable-length demonstrations. This avoids the brittleness of frame-index-based labeling and provides stable supervision even in tasks like T-shirt folding. Our reward model is robust to demonstration variability, generalizes to out-of-distribution scenarios, and improves downstream policy training. Building on it, we introduce Reward-Aligned Behavior Cloning (RA-BC), which filters and reweights demonstrations based on reward estimates. Experiments show that our method significantly outperforms baselines in both real-world rollouts and human validation. On T-shirt folding, we achieve 83% success from the flattened state and 67% from the crumpled state, compared to 8% and 0% with vanilla BC. Overall, our results highlight reward modeling as a scalable and annotation-efficient solution for long-horizon robotic manipulation. Project website: https://qianzhong-chen.github.io/sarm.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic ManipulationYuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li 等CVPR 2026 · 被引用 31 次
- ProgressLM: Towards Progress Reasoning in Vision-Language ModelsJianshu Zhang, Chengxuan Qian, Haosen Sun, Haoran Lu 等ACL 2026 · 被引用 7 次
- Recurrent Reasoning with Vision-Language Models for Estimating Long-Horizon Embodied Task ProgressYuelin Zhang, Sijie Cheng, Chen Li, Zongzhao Li 等CVPR 2026 · 被引用 6 次
- General Process Reward Modeling for Robotic Reinforcement LearningHuajie Tan, Sixiang Chen, Yijie Xu, Zixiao Wang 等CVPR 2026
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani 等ICML 2023 · 被引用 212 次
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement LearningXinyue Chen, Zijian Zhou, Zheng Wang, Che Wang 等NeurIPS 2020 · 被引用 146 次
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 被引用 135 次
- Video Prediction Models as Rewards for Reinforcement LearningAlejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain 等NeurIPS 2023 · 被引用 117 次
相关 Paper
- VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon ManipulationKuo-Han Hung, Pang-Chi Lo, Jia-Fong Yeh, Han-Yuan Hsu 等ICLR 2025
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang 等ICCV 2025 · 被引用 13 次
- Subtask-Aware Visual Reward Learning from Segmented DemonstrationsChangyeon Kim, Minho Heo, Doohyun Lee, Honglak Lee 等ICLR 2025
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic ManipulationNan Tang, Jing-Cheng Pang, Guanlin Li, Chao Qian 等ICML 2026 · 被引用 1 次
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji 等AAAI 2026 · 被引用 21 次
