ICML2026
Agentic Model Predictive Questioning Control in Visual Design
Kuang-Da Wang, Zhao Wang, Wei-Yao Wang, Yotaro Shimose, Jaechang Kim, Shingo Takamatsu
Abstract
Recent Large Language Model based approaches for clarifying visual design largely focus on selecting questions that better uncover user intent, but often overlooks the cognitive burden imposed on users, i.e., the effort required to interpret and answer these questions, which is crucial for effective human-agent interaction. In this paper, we propose Agentic Model Predictive Questioning Control (A-MPQC), a test-time framework that reduces proxy-estimated user interaction burden while improving visual design alignment by formulating multi-round clarification as trajectory optimization with receding-horizon replanning to revise its questioning strategy. In addition, we introduce lookahead question plans to reduce ambiguity early, and a lightweight respond-or-reject surrogate reward to steer questions toward lower user-burden formats (e.g., yes/no). Experiments on webpage and ad banner generation benchmarks show that A-MPQC not only generates designs better aligned with user intent, but also achieves lower user-interaction cost across diverse interaction baselines, including fixed-format strategies (e.g., multiple-choice and open-ended) and a retrieval-augmented baseline, without retraining. This paper sets a new perspective that explicitly formulates and optimizes the human cognitive burden jointly with final design alignment, opening new opportunities to advance human-agent interaction.