Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
John Joon Young Chung, Melissa Roemmele, Max Kreminski
Abstract
We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Visual Story-Writing: Writing by Manipulating Visual Representations of StoriesDamien Masson, Zixin Zhao, Fanny ChevalierUIST 2025 · 6 citations
- Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-EditingKexue Fu, Jingfei Huang, Long Ling, Sumin Hong et al.CHI 2026 · 3 citations
- PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XREsen K. Tütüncü, Qian Zhou, Frederik Brudy, George W. Fitzmaurice et al.CHI 2026 · 2 citations
- Texterial: A Text-as-Material Interaction Paradigm for LLM-Mediated WritingJocelyn J. Shen, Nicolai Marquardt, Hugo Romat, Ken Hinckley et al.CHI 2026 · 1 citation
- Vidmento: Creating Video Stories through Context-Aware Expansion with Generative VideoCatherine Yeh, Anh Truong, Mira Dontcheva, Bryan WangCHI 2026 · 1 citation
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- Think Then React: Towards Unconstrained Action-to-Reaction Motion GenerationWenhui Tan, Boyuan Li, Chuhao Jin, Wenbing Huang et al.ICLR 2025
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao et al.SIGGRAPH 2024 · 39 citations
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance GuidanceLei Li, Angela DaiICML 2026
- Mind the Gap: The Divergence Between Human and LLM-Generated TasksYi-Long Lu, Jiajun Song, Chunhui Zhang, Wei WangAAAI 2026
