Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning
Joseph Early, Tom Bewley, Christine Evers, Sarvapali D. Ramchurn
Abstract
We generalise the problem of reward modelling (RM) for reinforcement learning (RL) to handle non-Markovian rewards. Existing work assumes that human evaluators observe each step in a trajectory independently when providing feedback on agent behaviour. In this work, we remove this assumption, extending RM to capture temporal dependencies in human assessment of trajectories. We show how RM can be approached as a multiple instance learning (MIL) problem, where trajectories are treated as bags with return labels, and steps within the trajectories are instances with unseen reward labels. We go on to develop new MIL models that are able to capture the time dependencies in labelled trajectories. We demonstrate on a range of RL tasks that our novel MIL models can reconstruct reward functions to a high level of accuracy, and can be used to train high-performing agent policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Direct Preference-based Policy Optimization without Reward ModelingGaon An, Junhyeok Lee, Xingdong Zuo, Norio Kosaka et al.NeurIPS 2023 · 61 citations
- Inherently Interpretable Time Series Classification via Multiple Instance LearningJoseph Early, Gavin K. C. Cheung, Kurt Cutajar, Hanting Xie et al.ICLR 2024 · 29 citations
- STAR: Efficient Preference-based Reinforcement Learning via Dual RegularizationFengshuo Bai, Rui Zhao, Hongming Zhang, Sijia Cui et al.NeurIPS 2025 · 13 citations
- Decoding Global Preferences: Temporal and Cooperative Dependency Modeling in Multi-Agent Preference-Based Reinforcement LearningTianchen Zhu, Yue Qiu, Haoyi Zhou, Jianxin LiAAAI 2024 · 9 citations
Builds on4
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg et al.ICML 2020 · 81 citations
- Model Agnostic Interpretability for Multiple Instance LearningJoseph Early, Christine Evers, Sarvapali D. RamchurnICLR 2022 · 15 citations
Related papers
- Reward Learning through Ranking Mean Squared ErrorChaitanya Kharyal, Calarina Muslimani, Matthew TaylorICML 2026
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Reinforcement Learning with Trajectory FeedbackYonathan Efroni, Nadav Merlis, Shie MannorAAAI 2021 · 48 citations
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
- SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman RavindranAAAI 2026
