In-Trajectory Inverse Reinforcement Learning: Learn Incrementally Before an Ongoing Trajectory Terminates
Shicheng Liu, Minghui Zhu
摘要
Inverse reinforcement learning (IRL) aims to learn a reward function and a corresponding policy that best fit the demonstrated trajectories of an expert. However, current IRL works cannot learn incrementally from an ongoing trajectory because they have to wait to collect at least one complete trajectory to learn. To bridge the gap, this paper considers the problem of learning a reward function and a corresponding policy while observing the initial state-action pair of an ongoing trajectory and keeping updating the learned reward and policy when new state-action pairs of the ongoing trajectory are observed. We formulate this problem as an online bi-level optimization problem where the upper level dynamically adjusts the learned reward according to the newly observed state-action pairs with the help of a meta-regularization term, and the lower level learns the corresponding policy. We propose a novel algorithm to solve this problem and guarantee that the algorithm achieves sub-linear local regret . If the reward function is linear, we prove that the proposed algorithm achieves sub-linear regret . Experiments are used to validate the proposed algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement LearningHaochen Zhang, Zhong Zheng, Lingzhou XueNeurIPS 2025 · 被引用 3 次
- Explainable Reinforcement Learning from Human Feedback to Improve AlignmentShicheng Liu, Siyuan Xu, Wenjie Qiu, Hangfan Zhang 等NeurIPS 2025 · 被引用 2 次
- Gap-Dependent Bounds for Q-Learning using Reference-Advantage DecompositionZhong Zheng, Haochen Zhang, Lingzhou XueICLR 2025
- UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement LearningShicheng Liu, Minghui ZhuICLR 2025
- Meta-Reinforcement Learning with Adaptation from Human Feedback via Preference-Order-Preserving Task EmbeddingSiyuan Xu, Minghui ZhuICML 2025
它引用的顶会 Paper30
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max ProblemsJiawei Zhang, Peijun Xiao, Ruoyu Sun, Zhi-Quan LuoNeurIPS 2020 · 被引用 130 次
- Learning from Teaching Regularization: Generalizable Correlations Should be Easy to ImitateCan Jin, Tong Che, Hongwu Peng, Yiyuan Li 等NeurIPS 2024 · 被引用 67 次
- Maximum-Likelihood Inverse Reinforcement Learning with Finite-Time GuaranteesSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2022 · 被引用 60 次
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov 等AAAI 2024 · 被引用 60 次
相关 Paper
- Inverse Reinforcement Learning from a Gradient-based LearnerGiorgia Ramponi, Gianluca Drappo, Marcello RestelliNeurIPS 2020 · 被引用 16 次
- When Demonstrations meet Generative World Models: A Maximum Likelihood Framework for Offline Inverse Reinforcement LearningSiliang Zeng, Chenliang Li, Alfredo García, Mingyi HongNeurIPS 2023 · 被引用 33 次
- Non-Adversarial Inverse Reinforcement Learning via Successor Feature MatchingArnav Kumar Jain, Harley Wiltzer, Jesse Farebrother, Irina Rish 等ICLR 2025
- Inverse Reinforcement Learning in a Continuous State Space with Formal GuaranteesGregory Dexter, Kevin Bello, Jean HonorioNeurIPS 2021 · 被引用 9 次
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
