Foresight Optimization for Strategic Reasoning in Large Language Models
Jessie Wang, Jiawen Duan, Jian Wang, Kaitao Song, Chunpu Xu, Johnny K. W. Ho, Fenggang Yu, Johan F. Hoorn, Wenjie Li
摘要
Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-based LLMs to perform effective decision-making abilities in multiagent environments, due to the absence of explicit foresight modeling. To this end, strategic reasoning, the most fundamental capability to anticipate the counterpart's behaviors and foresee its possible future actions, has been introduced to alleviate the above issues. Strategic reasoning is fundamental to effective decision-making in multi-agent environments, yet existing reasoning enhancement methods for LLMs do not explicitly capture its foresight nature. In this work, we introduce Foresight Policy Optimization (FoPO) to enhance strategic reasoning in LLMs, which integrates opponent modeling principles into policy optimization, thereby enabling explicit consideration of both self-interest and counterpart influence. Specifically, we construct two curated datasets, namely Cooperative RSA and Competitive Taboo, equipped with well-designed rules and moderate difficulty to facilitate a systematic investigation of FoPO in a self-play framework. Our experiments demonstrate that FoPO significantly enhances strategic reasoning across LLMs of varying sizes and origins. Moreover, models trained with FoPO exhibit strong generalization to out-of-domain strategic scenarios, substantially outperforming standard LLM reasoning optimization baselines. 1 * Equal contribution. 1 https://github.com/wangjs9/ForesightOptim .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine 等ICML 2024 · 被引用 163 次
相关 Paper
- EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement LearningXiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu 等ACL 2025 · 被引用 7 次
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 被引用 10 次
- ReActR: Reasoning through Error-Activated Reflection for LLM Post-TrainingLina SunACL 2026
- CFPO: Counterfactual Policy Optimization for Multimodal ReasoningZhangyuan Yu, Wanran Sun, Guangjing Yang, Xiaohu Wu 等ICML 2026
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable ReasoningYuyang Ding, Chi Zhang, Juntao Li, Haibin Lin 等ICLR 2026 · 被引用 7 次
