Training a Generally Curious Agent
Fahim Tajwar, Yiding Jiang, Abitha Thankaraj, Sumaita Sadia Rahman, J. Zico Kolter, Jeff Schneider, Russ Salakhutdinov
摘要
Efficient exploration is essential for intelligent systems interacting with their environment, but existing language models often fall short in scenarios that require strategic information gathering. In this paper, we present PAPRIKA, a fine-tuning approach that enables language models to develop general decision-making capabilities that are not confined to particular environments. By training on synthetic interaction data from different tasks that require diverse strategies, PAPRIKA teaches models to explore and adapt their behavior on a new task based on environment feedback incontext without more gradient updates. Experimental results show that models fine-tuned with PAPRIKA can effectively transfer their learned decision-making capabilities to entirely unseen tasks without additional training. Unlike traditional training, our approach's primary bottleneck lies in sampling useful interaction data instead of model updates. To improve sample efficiency, we propose a curriculum learning strategy that prioritizes sampling trajectories from tasks with high learning potential. These results suggest a promising path towards AI systems that can autonomously solve novel sequential decisionmaking problems that require interactions with the external world.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling 等ICLR 2026 · 被引用 112 次
- Kevin: Multi-Turn RL for Generating CUDA KernelsCarlo Baronio, Pietro Marsella, Ben Pan, Simon Guo 等ICLR 2026 · 被引用 81 次
- Expanding the Capabilities of Reinforcement Learning via Text FeedbackYuda Song, Lili Chen, Fahim Tajwar, REMI MUNOS 等ICML 2026 · 被引用 41 次
- Meta-RL Induces Exploration in Language AgentsYulun Jiang, Liangze Jiang, Damien Teney, Michael Moor 等ICLR 2026 · 被引用 20 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- On the Modeling Capabilities of Large Language Models for Sequential Decision MakingMartin Klissarov, R. Devon Hjelm, Alexander T. Toshev, Bogdan MazoureICLR 2025
- Don't Just Fine-tune the Agent, Tune the EnvironmentSiyuan Lu, Zechuan Wang, Hongxuan Zhang, Qintong Wu 等ICLR 2026 · 被引用 13 次
- Active Example Selection for In-Context LearningYiming Zhang, Shi Feng, Chenhao TanEMNLP 2022 · 被引用 84 次
- Boosting Multi-Domain Fine-Tuning of Large Language Models through Evolving Interactions between SamplesXize Liang, Lin Yang, Jie Wang, Yiyang Lu 等ICML 2025
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
