TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference
Yulin Dou, Jiangming Liu
Abstract
Humans increasingly query Large Language Models (LLMs) to accomplish personal tasks according to their individual preferences. However, these preferences are often unconsciously veiled during conversation. To address this, LLMs have to elicit human preferences through multi-turn dialogue, where tasks are accomplished via iterative clarifying questions and final response generated by LLMs as effective questioners. Existing approaches based on self-taught reasoning have two limitations: 1) they struggle to avoid generating irrelevant questions and 2) the final responses to tasks are misled by the conversations. To overcome these limitations, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization. TO-GATE comprises two key components: a clarification resolver, which generates optimal questioning trajectories to produce effective elicitation questions, and a summarizer, which ensures task-aligned final responses. Experimental results show that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- When MOE Meets LLMs: Parameter Efficient Fine-tuning for Multi-task Medical ApplicationsQidong Liu, Xian Wu, Xiangyu Zhao, Yuanshao Zhu et al.SIGIR 2024 · 89 citations
- MultiGPrompt for Multi-Task Pre-Training and Prompting on GraphsXingtong Yu, Chang Zhou, Yuan Fang, Xinming ZhangWWW 2024 · 65 citations
Related papers
- Eliciting Human Preferences with Language ModelsBelinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob AndreasICLR 2025
- Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question AnsweringAlireza Salemi, Cheng Li, Mingyang Zhang, Qiaozhu Mei et al.WWW 2026 · 3 citations
- Modeling Future Conversation Turns to Teach LLMs to Ask Clarifying QuestionsMichael J. Q. Zhang, W. Bradley Knox, Eunsol ChoiICLR 2025
- Well Begun is Half Done: Generator-agnostic Knowledge Pre-Selection for Knowledge-Grounded DialogueLang Qin, Yao Zhang, Hongru Liang, Jun Wang et al.EMNLP 2023 · 1 citation
- PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based SamplingYongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang et al.EMNLP 2024 · 6 citations
