Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
Chenxing Wei, Hong Wang, Ying He, Zhongxiang Dai, Bo Jiang, Fei Yu, Yao Shu
Abstract
Test-time policy adaptation for multi-turn interactions (TPAM) is essential for aligning Large Language Models (LLMs) with dynamic user needs. However, existing paradigms typically treat adaptation as a single-axis problem by either purely refining instructions or solely updating weights. This bifurcated approach overlooks the fact that interaction failures arise from a coupled mixture of context ambiguity and model incapacity. To address this, we propose ROSA2, a framework that reformulates TPAM as a joint optimization problem over the heterogeneous space of Words and Weights. Within this framework, the semantic stream acts as a feedback normalizer that transforms noisy or ambiguous user feedback into actionable instructions, ensuring that parametric adaptation is performed on semantically clarified trajectories. Theoretically, we prove that this semantic pre-conditioning strictly reduces the required parameter shift for convergence. Empirically, ROSA2 consistently outperforms state-of-the-art baselines across diverse mathematical, general reasoning, and coding benchmarks. It achieves up to a 37.8% accuracy improvement while reducing average interaction turns by 40%, demonstrating that ROSA2 unlocks the true potential of parameter updates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46ad9a07-fe96-4d6b-be90-906070fe6cfdBuilds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
Related papers
- Test-Time Learning for Large Language ModelsJinwu Hu, Zitian Zhang, Guohao Chen, Xutao Wen et al.ICML 2025
- Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative GenerationHanwen Shen, Ting Ying, Jiajie Lu, Shanshan WangACL 2026 · 5 citations
- Talk2Code: A Multi-Turn Interaction Benchmark with Dual-Track Evaluation for Code GenerationWeibin Yang, Liangru Xie, Jieyun Cai, Yuxiang Yan et al.AAAI 2026
- Test-time Prompt InterventionChenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao et al.AAAI 2026 · 8 citations
- Parrot: Enhancing Multi-Turn Instruction Following for Large Language ModelsYuchong Sun, Che Liu, Kun Zhou, Jinwen Huang et al.ACL 2024
