Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedback
Mohd Hozaifa Khan, Ravi Kiran Sarvadevabhatla
Abstract
We introduce Sketchtopia, a large-scale dataset and AI framework designed to explore goal-driven, multimodal communication through asynchronous interactions in a Pictionary-inspired setup. Sketchtopia captures natural human interactions, including freehand sketches, open-ended guesses, and iconic feedback gestures, showcasing the complex dynamics of cooperative communication under constraints. It features over 20K gameplay sessions from 916 players, capturing 263K sketches, 10K erases, 56K guesses and 19.4K iconic feedbacks. We introduce multimodal foundational agents with capabilities for generative sketching, guess generation and asynchronous communication. Our dataset also includes 800 human-agent sessions for benchmarking the agents. We introduce novel metrics to characterize collaborative success, responsiveness to feedback and inter-agent asynchronous communication. Sketchtopia pushes the boundaries of multimodal AI, establishing a new benchmark for studying asynchronous, goal-oriented interactions between humans and AI agents. The dataset can be found at https://sketchtopia25.github.io/ Related Works Sketch datasets: Existing sketch datasets primarily focus on recognition [13, 40] , segmentation [22, 26, 36, 47, 49, 51] , and retrieval [53] , with labels assigned to the final sketch. In contrast, our dataset includes intermediate text "guess" annotations from gameplay, providing temporal grounding for sketches throughout the drawing process. Additionally, our dataset covers a wider array of categories, including abstract concepts (nouns, verbs, adjectives). Some works incorporate text guessing to simulate Pictionary [41] , but lack the complexity and richness generated by interacting agents and gameplay dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4b12e26-b9f7-4b2f-a531-b03f098bc169Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and TextChristopher Clark, Jordi Salvador, Dustin Schwenk, Derrick Bonafilia et al.EMNLP 2021 · 6 citations
- SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM InteractionZeyuan Huang, Cangjun Gao, Yaxian Shan, Haoxiang Hu et al.UIST 2025 · 8 citations
- DrawMon: A Distributed System for Detection of Atypical Sketch Content in Concurrent Pictionary GamesNikhil Bansal, Kartik Gupta, Kiruthika Kannan, Sivani Pentapati et al.ACM MM 2022 · 1 citation
- STICKERCONV: Generating Multimodal Empathetic Responses from ScratchYiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun et al.ACL 2024
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
