Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
John Joon Young Chung, Melissa Roemmele, Max Kreminski
摘要
We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and nuanced social interactions; Toyteller leverages these motions (1) to let users steer story text generation and (2) as a visual output format that accompanies story text. We enabled motion-steered text generation and text-steered motion generation by mapping motions and text onto a shared semantic space so that large language models and motion generation models can use it as a translational layer. Technical evaluations showed that Toyteller outperforms a competitive baseline, GPT-4o. Our user study identified that toy-playing helps express intentions difficult to verbalize. However, only motions could not express all user intentions, suggesting combining it with other modalities like language. We discuss the design space of toy-playing interactions and implications for technical HCI research on human-AI interaction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Visual Story-Writing: Writing by Manipulating Visual Representations of StoriesDamien Masson, Zixin Zhao, Fanny ChevalierUIST 2025 · 被引用 6 次
- Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-EditingKexue Fu, Jingfei Huang, Long Ling, Sumin Hong 等CHI 2026 · 被引用 3 次
- PlayWrite: A Multimodal System for AI Supported Narrative Co-Authoring Through Play in XREsen K. Tütüncü, Qian Zhou, Frederik Brudy, George W. Fitzmaurice 等CHI 2026 · 被引用 2 次
- Texterial: A Text-as-Material Interaction Paradigm for LLM-Mediated WritingJocelyn J. Shen, Nicolai Marquardt, Hugo Romat, Ken Hinckley 等CHI 2026 · 被引用 1 次
- Vidmento: Creating Video Stories through Context-Aware Expansion with Generative VideoCatherine Yeh, Anh Truong, Mira Dontcheva, Bryan WangCHI 2026 · 被引用 1 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
相关 Paper
- Think Then React: Towards Unconstrained Action-to-Reaction Motion GenerationWenhui Tan, Boyuan Li, Chuhao Jin, Wenbing Huang 等ICLR 2025
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture SynthesisZeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 等SIGGRAPH 2024 · 被引用 39 次
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated NarrationsQing Chang, Zhiming HuAAAI 2026
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance GuidanceLei Li, Angela DaiICML 2026
- Mind the Gap: The Divergence Between Human and LLM-Generated TasksYi-Long Lu, Jiajun Song, Chunhui Zhang, Wei WangAAAI 2026
