Closing the Expression Gap in LLM Instructions via Socratic Questioning
Jianwen Sun, Yukang Feng, Yifan Chang, Chuanhao Li, Zizhen Li, Jiaxin Ai, Fanrui Zhang, Yu Dai, Kaipeng Zhang
摘要
A fundamental bottleneck in human-AI collaboration is the "intention expression gap", the difficulty for humans to effectively convey complex, high-dimensional thoughts to AI. This challenge often traps users in inefficient trial-and-error loops and is exacerbated by the diverse expertise levels of users. We reframe this problem from passive instruction following to a Socratic collaboration paradigm, proposing an agent that actively probes for information to resolve its uncertainty about user intent. we name the proposed agent Nous, trained to acquire proficiency in this inquiry policy. The core mechanism of Nous is a training framework grounded in the first principles of information theory. Within this framework, we define the information gain from dialogue as an intrinsic reward signal, which is fundamentally equivalent to the reduction of Shannon entropy over a structured task space. This reward design enables us to avoid reliance on costly human preference annotations or external reward models. To validate our framework, we develop an automated simulation pipeline to generate a large-scale, preference-based dataset for the challenging task of scientific diagram generation. Comprehensive experiments, including ablations, subjective and objective evaluations, and tests across user expertise levels, demonstrate the effectiveness of our proposed framework. Nous achieves leading efficiency and output quality, while remaining robust to varying user expertise. In conclusion, our research provides a systematic methodology and a new perspective for addressing the issue of ambiguous intentions in complex human-machine collaboration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Can Large Language Models Serve as Rational Players in Game Theory? A Systematic AnalysisCaoyun Fan, Jindou Chen, Yaohui Jin, Hao HeAAAI 2024 · 被引用 123 次
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2022 · 被引用 95 次
相关 Paper
- MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt OptimizationJian Zhang, Zhangqi Wang, Haiping Zhu, Kangda Cheng 等AAAI 2026 · 被引用 9 次
- Uncertainty-Aware Clarification in LLM Agents with Information GainMengyi DENG, Zhiwei Li, Xin Li, Tingyu ZHU 等ICML 2026 · 被引用 1 次
- NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding ModificationWenshuo Zhang, Leixian Shen, Shuchang Xu, Jindu Wang 等UIST 2025 · 被引用 9 次
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun 等CHI 2026 · 被引用 2 次
- User-Aware Active Knowledge Acquisition for Emotional Support DialogueMufan Xu, Kehai Chen, Jiahao Hu, Xinchao Xu 等ICML 2026
