Agentic Model Predictive Questioning Control in Visual Design
Kuang-Da Wang, Zhao Wang, Wei-Yao Wang, Yotaro Shimose, Jaechang Kim, Shingo Takamatsu
Abstract
Recent Large Language Model based approaches for clarifying visual design largely focus on selecting questions that better uncover user intent, but often overlooks the cognitive burden imposed on users, i.e., the effort required to interpret and answer these questions, which is crucial for effective human-agent interaction. In this paper, we propose Agentic Model Predictive Questioning Control (A-MPQC), a test-time framework that reduces proxy-estimated user interaction burden while improving visual design alignment by formulating multi-round clarification as trajectory optimization with receding-horizon replanning to revise its questioning strategy. In addition, we introduce lookahead question plans to reduce ambiguity early, and a lightweight respond-or-reject surrogate reward to steer questions toward lower user-burden formats (e.g., yes/no). Experiments on webpage and ad banner generation benchmarks show that A-MPQC not only generates designs better aligned with user intent, but also achieves lower user-interaction cost across diverse interaction baselines, including fixed-format strategies (e.g., multiple-choice and open-ended) and a retrieval-augmented baseline, without retraining. This paper sets a new perspective that explicitly formulates and optimizes the human cognitive burden jointly with final design alignment, opening new opportunities to advance human-agent interaction. Our code is publicly available at https://github.com/sony/a_mpqc
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2aa6159-201c-49b2-9e60-615d378699fdBuilds on14
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
- ARGS: Alignment as Reward-Guided SearchMaxim Khanov, Jirayu Burapacheep, Yixuan LiICLR 2024 · 101 citations
- Aligning Large Language Models with Representation Editing: A Control PerspectiveLingkai Kong, Haorui Wang, Wenhao Mu, Yuanqi Du et al.NeurIPS 2024 · 80 citations
- Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye, Wei Xiong, Yuheng Zhang, Hanze Dong et al.NeurIPS 2024 · 60 citations
Related papers
- Proactive Agents for Multi-Turn Text-to-Image Generation Under UncertaintyMeera Hahn, Wenjun Zeng, Nithish Kannen, Rich Galt et al.ICML 2025
- Active Task Disambiguation with LLMsKasia Kobalczyk, Nicolás Astorga, Tennison Liu, Mihaela van der SchaarICLR 2025
- Uncertainty-Aware Clarification in LLM Agents with Information GainMengyi DENG, Zhiwei Li, Xin Li, Tingyu ZHU et al.ICML 2026 · 1 citation
- Prism: Towards Lowering User Cognitive Load in LLMs via Complex Intent UnderstandingZenghua Liao, Jinzhi Liao, Xiang ZhaoWWW 2026
- Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User InterfaceWenyue Hua, Mengting Wan, Jagannath Shashank Subramanya Sai Vadrevu, Ryan Nadel et al.ICLR 2025
