Targeted Data Acquisition for Evolving Negotiation Agents
Minae Kwon, Siddharth Karamcheti, Mariano-Florentino Cuellar, Dorsa Sadigh
摘要
Successful negotiators must learn how to balance optimizing for self-interest and cooperation. Yet current artificial negotiation agents often heavily depend on the quality of the static datasets they were trained on, limiting their capacity to fashion an adaptive response balancing self-interest and cooperation. For this reason, we find that these agents can achieve either high utility or cooperation, but not both. To address this, we introduce a targeted data acquisition framework where we guide the exploration of a reinforcement learning agent using annotations from an expert oracle. The guided exploration incentivizes the learning agent to go beyond its static dataset and develop new negotiation strategies. We show that this enables our agents to obtain higher-reward and more Pareto-optimal solutions when negotiating with both simulated and human partners compared to standard supervised learning and reinforcement learning methods. This trend additionally holds when comparing agents using our targeted data acquisition framework to variants of agents trained with a mix of supervised learning and reinforcement learning, or to agents using tailored reward functions that explicitly optimize for utility and Pareto-optimality. Code can be found here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
- On the Critical Role of Conventions in Adaptive Human-AI CollaborationAndy Shih, Arjun Sawhney, Jovana Kondic, Stefano Ermon 等ICLR 2021 · 被引用 46 次
- On the interaction between supervision and self-play in emergent communicationRyan Lowe, Abhinav Gupta, Jakob N. Foerster, Douwe Kiela 等ICLR 2020 · 被引用 30 次
相关 Paper
- ACE: A LLM-based Negotiation Coaching SystemRyan Shea, Aymen Kallala, Xin Liu, Michael W. Morris 等EMNLP 2024 · 被引用 6 次
- Be Selfish, But Wisely: Investigating the Impact of Agent Personality in Mixed-Motive Human-Agent InteractionsKushal Chawla, Ian Wu, Yu Rong, Gale M. Lucas 等EMNLP 2023 · 被引用 3 次
- GENTEEL-NEGOTIATOR: LLM-Enhanced Mixture-of-Expert-Based Reinforcement Learning Approach for Polite Negotiation DialoguePriyanshu Priya, Rishikant Chigrupaatii, Mauajama Firdaus, Asif EkbalAAAI 2025 · 被引用 14 次
- GREIL-Crowds: Crowd Simulation with Deep Reinforcement Learning and ExamplesPanayiotis Charalambous, Julien Pettré, Vassilis Vassiliades, Yiorgos Chrysanthou 等SIGGRAPH 2023 · 被引用 48 次
- TGRL: An Algorithm for Teacher Guided Reinforcement LearningIdan Shenfeld, Zhang-Wei Hong, Aviv Tamar, Pulkit AgrawalICML 2023 · 被引用 22 次
