Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
Jian Wang, Chak Tou Leong, Jiashuo Wang, Dongding Lin, Wenjie Li, Xiaoyong Wei
摘要
Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents. Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role disparities between two speakers and the multi-round interactive process that dialogues ought to be. Such a manner often leads to unsatisfactory chat consistency for the built agent. In this work, we emphasize the interactive, communicative nature of dialogue and argue that it is more feasible to model the speaker roles of agent and user separately, enabling the agent to adhere to its role consistently. With this in mind, we propose an efficient Multi-round Interactive Dialogue Tuning (MIDI-Tuning) framework 1 . It models the agent and user individually with two adapters built upon large language models. The adapters make use of respective utterances round by round in alternating order and they are tuned via a round-level memory caching mechanism. Extensive experiments demonstrate that, our framework performs superior to traditional finetuning and harbors the tremendous potential for improving dialogue consistency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from ScratchJiawei Chen, Xinyan Guan, Qianhao Yuan, Guozhao Mo 等EMNLP 2025
- Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational RecommendationDongding Lin, Jian Wang, Yongqi Li, Wenjie LiACL 2026
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
相关 Paper
- Consistently Simulating Human Personas with Multi-Turn Reinforcement LearningMarwa Abdulhai, Ryan Cheng, Donovan Clay, Tim Althoff 等NeurIPS 2025 · 被引用 51 次
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-RolesJing Han, Binwei Yan, Tianyu Guo, Zheyuan Bai 等ICML 2025
- "In-Dialogues We Learn": Towards Personalized Dialogue Without Pre-defined Profiles through In-Dialogue LearningChuanqi Cheng, Quan Tu, Wei Wu, Shuo Shang 等EMNLP 2024 · 被引用 3 次
- Learning to Know Myself: A Coarse-to-Fine Persona-Aware Training Framework for Personalized Dialogue GenerationYunpeng Li, Yue Hu, Yajing Sun, Luxi Xing 等AAAI 2023 · 被引用 9 次
- BoB: BERT Over BERT for Training Persona-based Dialogue Models from Limited Personalized DataHaoyu Song, Yan Wang, Kaiyan Zhang, Wei-Nan Zhang 等ACL 2021
