Generating Illustrated Instructions
Sachit Menon, Ishan Misra, Rohit Girdhar
Abstract
We introduce a new task of generating “Illustrated Instructions ”, i.e. visual instructions customized to a user's needs. We identify desiderata unique to this task, and formalize it through a suite of automatic and human evaluation metrics, designed to measure the validity, consistency, and efficacy of the generations. We combine the power of large language models (LLMs) together with strong text-to-image generation diffusion models to propose a simple approach called StackedDiffusion, which generates such illustrated instructions given text as input. The resulting model strongly outperforms baseline approaches and state-of-the-art multimodal LLMs; and in 30% of cases, users even prefer it to human-generated articles. Most notably, it enables various new and exciting applications far beyond what static articles on the web can provide, such as personalized instructions complete with intermediate steps and pictures in response to a user's individual situation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 077bf20c-3178-4fe7-874a-c9d4f2d08c42Cited by top-tier papers5
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 1 citation
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami et al.ICCV 2025 · 1 citation
- ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual InstructionsTomás Soucek, Prajwal Gatti, Michael Wray, Ivan Laptev et al.CVPR 2025
- Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflectionYucheng Suo, Fan Ma, Kaixin Shen, Linchao Zhu et al.ICLR 2025
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationRuoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao et al.ACM MM 2025
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Generating Coherent Sequences of Visual Illustrations for Real-World Manual TasksJoão Bordalo, Vasco Ramos, Rodrigo Valerio, Diogo Glória-Silva et al.ACL 2024 · 1 citation
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang et al.ICLR 2024 · 173 citations
- Customization Assistant for Text-to-image GenerationYufan Zhou, Ruiyi Zhang, Jiuxiang Gu, Tong SunCVPR 2024 · 10 citations
- Instruct-Imagen: Image Generation with Multi-modal InstructionHexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su, Wenhu Chen et al.CVPR 2024
- Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision ModelsAda Yi Zhao, Aditya Gunturu, Ellen Yi-Luen Do, Ryo SuzukiUIST 2025 · 12 citations
