CollabLLM: From Passive Responders to Active Collaborators
Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, Jianfeng Gao
摘要
Website: aka.ms/CollabLLM ④ Multiturn-aware Reward ② Response 𝒚 Real-world or Simulated User Policy 𝝅 𝜽 𝒚 𝒙 ③ Collaborative Simulation Forward Sampling Reward Computation #1 #2 #3 ① Context state (𝒙) I need to write about how optimism can improve our well-being. To get us started, what kind of tone are you aiming for? Online generation RL finetuning #1 #2 #3 … … (𝒙, 𝒚) Extrinsic Reward e.g., Performance Intrinsic Reward Interactivity Efficiency Figure 1: COLLABLLM Framework: Given a context 1 , the model generates a response 2 to maximize long-term collaboration gains, termed Multiturn-aware Rewards (MR). During training, MRs are estimated via 3 collaborative simulation, which forward-samples conversations with simulated users. Finally, 4 reinforcement fine-tuning is applied using the MRs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent CollaborationYijia Shao, Vinay Samuel, Yucheng Jiang, John Yang 等ICLR 2026 · 被引用 57 次
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao 等ICLR 2026 · 被引用 24 次
- Can LLM Agents Simulate Multi-Turn Human Behavior? Evidence from Real Online Customer Behavior DataYuxuan Lu, Jing Huang, Yan Han, Bingsheng Yao 等ACL 2026 · 被引用 17 次
- Value of Information: A Framework for Human-Agent CommunicationYijiang River Dong, Tiancheng Hu, Zheng Hui, Caiqi Zhang 等ACL 2026 · 被引用 8 次
- GUIDE: A Benchmark for Understanding and Assisting Users in Open-Ended GUI TasksSaelyne Yang, Jaesang Yu, Yi-Hao Peng, Kevin Qinghong Lin 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett 等WWW 2020 · 被引用 238 次
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng 等ICLR 2024 · 被引用 86 次
相关 Paper
- LLM Collaboration with Multi-Agent Reinforcement LearningShuo Liu, Zeyu Liang, Xueguang Lyu, Christopher AmatoAAAI 2026
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement LearningChanwoo Park, Seungju Han, Xingzhi Guo, Asuman E. Ozdaglar 等ACL 2025 · 被引用 63 次
- InfoPO: Information-Driven Policy Optimization for User-Centric AgentsFanqi Kong, Jiayi Zhang, Mingyi Deng, Chenglin Wu 等ICML 2026
- Implicit Turn-Wise Policy Optimization for Proactive User-LLM InteractionHaoyu Wang, Yuxin Chen, Liang Luo, Buyun Zhang 等ICML 2026 · 被引用 3 次
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu 等NeurIPS 2024 · 被引用 54 次
