TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference
Yulin Dou, Jiangming Liu
摘要
Humans increasingly query Large Language Models (LLMs) to accomplish personal tasks according to their individual preferences. However, these preferences are often unconsciously veiled during conversation. To address this, LLMs have to elicit human preferences through multi-turn dialogue, where tasks are accomplished via iterative clarifying questions and final response generated by LLMs as effective questioners. Existing approaches based on self-taught reasoning have two limitations: 1) they struggle to avoid generating irrelevant questions and 2) the final responses to tasks are misled by the conversations. To overcome these limitations, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization. TO-GATE comprises two key components: a clarification resolver, which generates optimal questioning trajectories to produce effective elicitation questions, and a summarizer, which ensures task-aligned final responses. Experimental results show that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical ApplicationsQidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu 等SIGIR 2024 · 被引用 89 次
- MultiGPrompt for Multi-Task Pre-Training and Prompting on GraphsXingtong Yu, Chang Zhou, Yuan Fang, Xinming ZhangWWW 2024 · 被引用 65 次
相关 Paper
- Eliciting Human Preferences with Language ModelsBelinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob AndreasICLR 2025
- Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question AnsweringAlireza Salemi, Cheng Li, Mingyang Zhang, Qiaozhu Mei 等WWW 2026 · 被引用 3 次
- Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying QuestionsMichael J. Q. Zhang, W. Bradley Knox, Eunsol ChoiICLR 2025
- Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded DialogueLang Qin, Yao Zhang, Hongru Liang, Jun Wang 等EMNLP 2023 · 被引用 1 次
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 等EMNLP 2024 · 被引用 6 次
