Lune

ICML2025Top-tier venue

CollabLLM: From Passive Responders to Active Collaborators

Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng, Gavin Li, Yao Dou, Weixin Cai, James Zou, Jure Leskovec, Jianfeng Gao

2025Year
17Top-tier citations

Abstract

Website: aka.ms/CollabLLM โ‘ฃ Multiturn-aware Reward โ‘ก Response ๐’š Real-world or Simulated User Policy ๐… ๐œฝ ๐’š ๐’™ โ‘ข Collaborative Simulation Forward Sampling Reward Computation #1 #2 #3 โ‘  Context state (๐’™) I need to write about how optimism can improve our well-being. To get us started, what kind of tone are you aiming for? Online generation RL finetuning #1 #2 #3 โ€ฆ โ€ฆ (๐’™, ๐’š) Extrinsic Reward e.g., Performance Intrinsic Reward Interactivity Efficiency Figure 1: COLLABLLM Framework: Given a context 1 , the model generates a response 2 to maximize long-term collaboration gains, termed Multiturn-aware Rewards (MR). During training, MRs are estimated via 3 collaborative simulation, which forward-samples conversations with simulated users. Finally, 4 reinforcement fine-tuning is applied using the MRs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b6828db0-e10e-48ec-b9b8-b07e184dfc23

Cited by top-tier papers17

Ask how each one uses it

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines