SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Zeyuan Huang, Cangjun Gao, Yaxian Shan, Haoxiang Hu, Qingkun Li, Xiaoming Deng, Cuixia Ma, Yu-Kun Lai, Yong-Jin Liu, Feng Tian, Guozhong Dai, Hongan Wang
摘要
Human interaction with large language models (LLMs) is typically confined to text or image interfaces. Sketches offer a powerful medium for articulating creative ideas and user intentions, yet their potential remains underexplored. We propose SketchGPT, a novel interaction paradigm that integrates sketch and speech input directly over the system interface, facilitating open-ended, context-aware communication with LLMs. By leveraging the complementary strengths of multimodal inputs, expressions are enriched with semantic scope while maintaining efficiency. Interpreting user intentions across diverse contexts and modalities remains a key challenge. To address this, we developed a prototype based on a multi-agent framework that infers user intentions within context and generates executable context-sensitive and toolkit-aware feedback. Using Chain-of-Thought techniques for temporal and semantic alignment, the system understands multimodal intentions and performs operations following human-in-the-loop confirmation to ensure reliability. User studies demonstrate that SketchGPT significantly outperforms unimodal manipulation approaches, offering more intuitive and effective means to interact with LLMs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun 等CHI 2026 · 被引用 2 次
- SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation GenerationBoyu Li, Lin-Ping Yuan, Zeyu Wang, Hongbo FuCHI 2026 · 被引用 1 次
- JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric ModelingRunlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu 等CHI 2026 · 被引用 1 次
- AutoGraph: Enabling Visual Context via Graph Alignment in Open Domain Multi-Modal Dialogue GenerationDeji Zhao, Donghong Han, Ye Yuan, Bo Ning 等ACM MM 2024 · 被引用 3 次
- SketchAgent: Language-Driven Sequential Sketch GenerationYael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao 等CVPR 2025
