VORTEX: Aligning Task Utility and Human Preferences Through LLM-Guided Reward Shaping
Guojun Xiong, Milind Tambe
摘要
In social impact optimization, AI decision systems often rely on solvers that optimize well-calibrated mathematical objectives. However, these solvers cannot directly accommodate evolving human preferences, typically expressed in natural language rather than formal constraints. Recent approaches address this by using large language models (LLMs) to generate new reward functions from preference descriptions. While flexible, they risk sacrificing the system's core utility guarantees. In this paper, we propose VORTEX, a language-guided reward shaping framework that preserves established optimization goals while adaptively incorporating human feedback. By formalizing the problem as multi-objective optimization, we use LLMs to iteratively generate shaping rewards based on verbal reinforcement and text-gradient prompt updates. This allows stakeholders to steer decision behavior via natural language without modifying solvers or specifying trade-off weights. We provide theoretical guarantees that VORTEX converges to Pareto-optimal trade-offs between utility and preference satisfaction. Empirical results in real-world allocation tasks demonstrate that VORTEX outperforms baselines in satisfying human-aligned coverage goals while maintaining high task performance. This work introduces a practical and theoretically grounded paradigm for human-AI collaborative optimization guided by natural language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingYangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng 等NeurIPS 2024 · 被引用 197 次
- A Decision-Language Model (DLM) for Dynamic Restless Multi-Armed Bandit Tasks in Public HealthNikhil Behari, Edwin Zhang, Yunfan Zhao, Aparna Taneja 等NeurIPS 2024 · 被引用 39 次
- Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless BanditsGuojun Xiong, Jian Li, Rahul SinghAAAI 2022 · 被引用 23 次
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
- Learning Infinite-Horizon Average-Reward Restless Multi-Action Bandits via Index AwarenessGuojun Xiong, Shufan Wang, Jian LiNeurIPS 2022 · 被引用 21 次
相关 Paper
- REvolve: Reward Evolution with Large Language Models using Human FeedbackRishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi 等ICLR 2025
- R*: Efficient Reward Design via Reward Structure Evolution and Parameter Alignment Optimization with Large Language ModelsPengyi Li, Jianye Hao, Hongyao Tang, Yifu Yuan 等ICML 2025
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language ModelsChengao Li, Hanyu Zhang, Yunkun Xu, Hongyan Xue 等ACL 2025 · 被引用 13 次
- Language Instructed Reinforcement Learning for Human-AI CoordinationHengyuan Hu, Dorsa SadighICML 2023 · 被引用 90 次
- Strategic Planning: A Top-Down Approach to Option GenerationMax Ruiz Luyten, Antonin Berthon, Mihaela van der SchaarICML 2025
