SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction
Zeyuan Huang, Cangjun Gao, Yaxian Shan, Haoxiang Hu, Qingkun Li, Xiaoming Deng, Cuixia Ma, Yu-Kun Lai, Yong-Jin Liu, Feng Tian, Guozhong Dai, Hongan Wang
Abstract
Human interaction with large language models (LLMs) is typically confined to text or image interfaces. Sketches offer a powerful medium for articulating creative ideas and user intentions, yet their potential remains underexplored. We propose SketchGPT, a novel interaction paradigm that integrates sketch and speech input directly over the system interface, facilitating open-ended, context-aware communication with LLMs. By leveraging the complementary strengths of multimodal inputs, expressions are enriched with semantic scope while maintaining efficiency. Interpreting user intentions across diverse contexts and modalities remains a key challenge. To address this, we developed a prototype based on a multi-agent framework that infers user intentions within context and generates executable context-sensitive and toolkit-aware feedback. Using Chain-of-Thought techniques for temporal and semantic alignment, the system understands multimodal intentions and performs operations following human-in-the-loop confirmation to ensure reliability. User studies demonstrate that SketchGPT significantly outperforms unimodal manipulation approaches, offering more intuitive and effective means to interact with LLMs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 30799e62-f738-4367-ba91-fbb6b0d762a2Cited by top-tier papers1
Ask how each one uses itRelated papers
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun et al.CHI 2026 · 2 citations
- SketchDynamics: Exploring Free-Form Sketches for Dynamic Intent Expression in Animation GenerationBoyu Li, Lin-Ping Yuan, Zeyu Wang, Hongbo FuCHI 2026 · 1 citation
- JustShape: Exploring Co-Speech Gestures for Multimodal LLM-Powered 3D Parametric ModelingRunlin Duan, Yuzhao Chen, Yichen Hu, Ziyi Liu et al.CHI 2026 · 1 citation
- AutoGraph: Enabling Visual Context via Graph Alignment in Open Domain Multi-Modal Dialogue GenerationDeji Zhao, Donghong Han, Ye Yuan, Bo Ning et al.ACM MM 2024 · 3 citations
- SketchAgent: Language-Driven Sequential Sketch GenerationYael Vinker, Tamar Rott Shaham, Kristine Zheng, Alex Zhao et al.CVPR 2025
