Reward Design for Justifiable Sequential Decision-Making
Aleksa Sukovic, Goran Radanovic
Abstract
Equipping agents with the capacity to justify made decisions using supporting evidence represents a cornerstone of accountable decision-making. Furthermore, ensuring that justifications are in line with human expectations and societal norms is vital, especially in high-stakes situations such as healthcare. In this work, we propose the use of a debate-based reward model for reinforcement learning agents, where the outcome of a zero-sum debate game quantifies the justifiability of a decision in a particular state. This reward model is then used to train a justifiable policy, whose decisions can be more easily corroborated with supporting evidence. In the debate game, two argumentative agents take turns providing supporting evidence for two competing decisions. Given the proposed evidence, a proxy of a human judge evaluates which decision is better justified. We demonstrate the potential of our approach in learning policies for prescribing and justifying treatment decisions of septic patients. We show that augmenting the reward with the feedback signal generated by the debate-based reward model yields policies highly favored by the judge when compared to the policy obtained solely from the environment rewards, while hardly sacrificing any performance. Moreover, in terms of the overall performance and justifiability of trained policies, the debate-based feedback is comparable to the feedback obtained from an ideal judge proxy that evaluates decisions using the full information encoded in the state. This suggests that the debate game outputs key information contained in states that is most relevant for evaluating decisions, which in turn substantiates the practicality of combining our approach with human-in-the-loop evaluations. Lastly, we showcase that agents trained via multi-agent debate learn to propose evidence that is resilient to refutations and closely aligns with human preferences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 195c090b-1af5-49c5-b939-e3355fa7074bBuilds on6
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 60 citations
- Reasoning on Knowledge Graphs with Debate DynamicsMarcel Hildebrandt, Jorge Andres Quintero Serna, Yunpu Ma, Martin Ringsquandl et al.AAAI 2020 · 59 citations
- Contextual Games: Multi-Agent Learning with Side InformationPier Giuseppe Sessa, Ilija Bogunovic, Andreas Krause, Maryam KamgarpourNeurIPS 2020 · 26 citations
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 21 citations
Related papers
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math ReasoningDan Qiao, Binbin Chen, Fengyu Cai, Jianlong Chen et al.ICML 2026 · 3 citations
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan et al.AAAI 2026
- iMAD: Intelligent Multi-Agent Debate for Efficient and Accurate LLM InferenceWei Fan, JinYi Yoon, Bo JiAAAI 2026 · 5 citations
- Debate2Create: Robot Co-design via Multi-Agent LLM DebateKevin Qiu, Marek CyganICML 2026
- Online Decision MediationDaniel Jarrett, Alihan Hüyük, Mihaela van der SchaarNeurIPS 2022 · 5 citations
