Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning
Joseph Early, Tom Bewley, Christine Evers, Sarvapali D. Ramchurn
摘要
We generalise the problem of reward modelling (RM) for reinforcement learning (RL) to handle non-Markovian rewards. Existing work assumes that human evaluators observe each step in a trajectory independently when providing feedback on agent behaviour. In this work, we remove this assumption, extending RM to capture temporal dependencies in human assessment of trajectories. We show how RM can be approached as a multiple instance learning (MIL) problem, where trajectories are treated as bags with return labels, and steps within the trajectories are instances with unseen reward labels. We go on to develop new MIL models that are able to capture the time dependencies in labelled trajectories. We demonstrate on a range of RL tasks that our novel MIL models can reconstruct reward functions to a high level of accuracy, and can be used to train high-performing agent policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka 等NeurIPS 2023 · 被引用 61 次
- Inherently Interpretable Time Series Classification via Multiple Instance LearningJoseph Early, Gavin K. C. Cheung, Kurt Cutajar, Hanting Xie 等ICLR 2024 · 被引用 29 次
- STAR: Efficient Preference-based Reinforcement Learning via Dual RegularizationFengshuo Bai, Rui Zhao, Hongming Zhang, Sijia Cui 等NeurIPS 2025 · 被引用 13 次
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 被引用 9 次
它引用的顶会 Paper4
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang 等NeurIPS 2021 · 被引用 1,163 次
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 被引用 219 次
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Model Agnostic Interpretability for Multiple Instance LearningJoseph Early, Christine Evers, Sarvapali D. RamchurnICLR 2022 · 被引用 15 次
相关 Paper
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 被引用 48 次
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
- SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman RavindranAAAI 2026
