THEMES: An Offline Apprenticeship Learning Framework for Evolving Reward Functions
Xi Yang, Md. Mirajul Islam, Ge Gao, Min Chi
摘要
Apprenticeship learning (AL) aims to induce decision-making policies by observing and imitating expert demonstrations.Existing AL approaches typically rely on online interactions and assume that the demonstrations follow a single reward function.Nevertheless, in real-world human-centric applications, policies are usually learned in an offline setting, with the demonstrations driven by multiple reward functions that evolve over time.To address these challenges, we introduce a novel AL framework: Time-aware Hierarchical EM Energy-based Sub-trajectory (THEMES) clustering.We evaluate the effectiveness of THEMES in two challenging human-centric domains -healthcare and education.Our experimental results across multiple datasets demonstrate that THEMES can accurately induce policies, outperforming competitive baselines and ablations, demonstrating its potential for tackling a broad range of complex, real-world human-centric tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Strictly Batch Imitation Learning by Energy-based Distribution MatchingDaniel Jarrett, Ioana Bica, Mihaela van der SchaarNeurIPS 2020 · 被引用 74 次
- On Trajectory Augmentations for Off-Policy EvaluationGe Gao, Qitong Gao, Xi Yang, Song Ju 等ICLR 2024 · 被引用 5 次
- Noise-conditioned Energy-based Annealed Rewards (NEAR): A Generative Framework for Imitation Learning from ObservationAnish Abhijit Diwan, Julen Urain, Jens Kober, Jan PetersICLR 2025
- Interpretable and Personalized Apprenticeship Scheduling: Learning Interpretable Scheduling Policies from Heterogeneous User DemonstrationsRohan R. Paleja, Andrew Silva, Letian Chen, Matthew C. GombolayNeurIPS 2020 · 被引用 41 次
- Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of ExpertsOnur Celik, Aleksandar Taranovic, Gerhard NeumannICML 2024 · 被引用 19 次
