Select to Perfect: Imitating desired behavior from large multi-agent data
Tim Franzmeyer, Edith Elkind, Philip Torr, Jakob Nicolaus Foerster, João F. Henriques
摘要
AI agents are commonly trained with large datasets of demonstrations of human behavior. However, not all behaviors are equally safe or desirable. Desired characteristics for an AI agent can be expressed by assigning desirability scores, which we assume are not assigned to individual behaviors but to collective trajectories. For example, in a dataset of vehicle interactions, these scores might relate to the number of incidents that occurred. We first assess the effect of each individual agent's behavior on the collective desirability score, e.g., assessing how likely an agent is to cause incidents. This allows us to selectively imitate agents with a positive effect, e.g., only imitating agents that are unlikely to cause incidents. To enable this, we propose the concept of an agent's Exchange Value, which quantifies an individual agent's contribution to the collective desirability score. The Exchange Value is the expected change in desirability score when substituting the agent for a randomly selected agent. We propose additional methods for estimating Exchange Values from real-world datasets, enabling us to learn desired imitation policies that outperform relevant baselines. The project website can be found at https://tinyurl.com/select-to-perfect .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Strategy Representation for Imitation Learning in Multi-Agent GamesShiqi Lei, Kanghoon Lee, Linjing Li, Jinkyoo ParkAAAI 2025 · 被引用 1 次
- Offline Multi-Agent Reinforcement Learning via Sequential Score DecompositionDan Qiao, Wenhao Li, Shanchao Yang, Hongyuan Zha 等ICML 2026
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 被引用 159 次
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler 等NeurIPS 2020 · 被引用 124 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
相关 Paper
- SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred TrajectoriesReturaj Burnwal, Nirav Pravinbhai Bhatt, Balaraman RavindranAAAI 2026
- Multi-Agent Interactions Modeling with Correlated PoliciesMinghuan Liu, Ming Zhou, Weinan Zhang, Yuzheng Zhuang 等ICLR 2020 · 被引用 22 次
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
- Limited Preference Aided Imitation Learning from Imperfect DemonstrationsXingchen Cao, Fan-Ming Luo, Junyin Ye, Tian Xu 等ICML 2024 · 被引用 6 次
- DualCOIL: Offline Imitation Learning from Contrasting DemonstrationsHuy Hoang, Tien Mai, Pradeep Varakantham, Tanvi VermaICML 2026
