Neural Dynamic Policies for End-to-End Sensorimotor Learning
Shikhar Bahl, Mustafa Mukadam, Abhinav Gupta, Deepak Pathak
摘要
The current dominant paradigm in sensorimotor control, whether imitation or reinforcement learning, is to train policies directly in raw action spaces such as torque, joint angle, or end-effector position. This forces the agent to make decisions individually at each timestep in training, and hence, limits the scalability to continuous, high-dimensional, and long-horizon tasks. In contrast, research in classical robotics has, for a long time, exploited dynamical systems as a policy representation to learn robot behaviors via demonstrations. These techniques, however, lack the flexibility and generalizability provided by deep learning or reinforcement learning and have remained under-explored in such settings. In this work, we begin to close this gap and embed the structure of a dynamical system into deep neural network-based policies by reparameterizing action spaces via second-order differential equations. We propose Neural Dynamic Policies (NDPs) that make predictions in trajectory distribution space as opposed to prior policy learning methods where actions represent the raw control space. The embedded structure allows end-to-end policy learning for both reinforcement and imitation learning setups. We show that NDPs outperform the prior state-of-the-art in terms of either efficiency or performance across several robotic control tasks for both imitation and reinforcement learning setups. Project video and code are available at: https://shikharbahl.github.io/neural-dynamic-policies/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Accelerating Robotic Reinforcement Learning via Parameterized Action PrimitivesMurtaza Dalal, Deepak Pathak, Ruslan SalakhutdinovNeurIPS 2021 · 被引用 121 次
- Hierarchical Few-Shot Imitation with Skill Transition ModelsKourosh Hakhamaneshi, Ruihan Zhao, Albert Zhan, Pieter Abbeel 等ICLR 2022 · 被引用 51 次
- Wasserstein Gradient Flows for Optimizing Gaussian Mixture PoliciesHanna Ziesche, Leonel RozoNeurIPS 2023 · 被引用 12 次
- Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement LearningGe Li, Hongyi Zhou, Dominik Roth, Serge Thilges 等ICLR 2024 · 被引用 11 次
- Learning Routines for Effective Off-Policy Reinforcement LearningEdoardo Cetin, Oya ÇeliktutanICML 2021 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation LearningWeikang Wan, Ziyu Wang, Yufei Wang, Zackory Erickson 等NeurIPS 2024
- Causal Navigation by Continuous-time Neural NetworksCharles Vorbach, Ramin M. Hasani, Alexander Amini, Mathias Lechner 等NeurIPS 2021 · 被引用 64 次
- Demystifying Action Space Design for Robotic Manipulation PoliciesYuchun Feng, Jinliang Zheng, Tianyi Tan, Dongxiu Liu 等ICML 2026 · 被引用 6 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action RepresentationBoyan Li, Hongyao Tang, Yan Zheng, Jianye Hao 等ICLR 2022 · 被引用 79 次
