Brickify: Enabling Expressive Design Intent Specification through Direct Manipulation on Design Tokens
Xinyu Shi, Yinghou Wang, Ryan A. Rossi, Jian Zhao
Abstract
Expressing design intent using natural language prompts requires designers to verbalize the ambiguous visual details concisely, which can be challenging or even impossible. To address this, we introduce Brickify, a visual-centric interaction paradigm — expressing design intent through direct manipulation on design tokens. Brickify extracts visual elements (e.g., subject, style, and color) from reference images and converts them into interactive and reusable design tokens that can be directly manipulated (e.g., resize, group, link, etc.) to form the visual lexicon. The lexicon reflects users’ intent for both what visual elements are desired and how to construct them into a whole. We developed Brickify to demonstrate how AI models can interpret and execute the visual lexicon through an end-to-end pipeline. In a user study, experienced designers found Brickify more efficient and intuitive than text-based prompts, allowing them to describe visual details, explore alternatives, and refine complex designs with greater ease and control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 769c5086-4d6a-46b4-93f0-25c1ed389f7dCited by top-tier papers9
- StoryEnsemble: Enabling Dynamic Exploration & Iteration in the Design Process with AI and Forward-Backward PropagationSangho Suh, Michael Lai, Kevin Pu, Steven P. Dow et al.UIST 2025 · 5 citations
- DataWink: Reusing and Adapting SVG-Based Visualization Examples with Large Multimodal ModelsLiwenhan Xie, Yanna Lin, Can Liu, Huamin Qu et al.IEEE VIS 2025 · 3 citations
- DesignTrace: Exploring, Iterating and Tracking Design Alternatives with GenAIXiaohan Peng, Debanjana Haldar, Wendy E. Mackay, Janin KochCHI 2026 · 3 citations
- Collaposer: Transforming Photo Collections into Visual Assets for Storytelling with CollagesJiayi Zhou, Liwenhan Xie, Jiaju Ma, Zheng Wei et al.CHI 2026 · 3 citations
- Vistoria: A Multimodal System to Support Fictional Story Writing through Instrumental Image-Text Co-EditingKexue Fu, Jingfei Huang, Long Ling, Sumin Hong et al.CHI 2026 · 3 citations
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore et al.UIST 2023 · 179 citations
- Bridging Gulfs in UI Generation through Semantic GuidanceSeokhyeon Park, Soohyun Lee, Eugene Choi, Hyunwoo Kim et al.CHI 2026 · 3 citations
- WorldSmith: Iterative and Expressive Prompting for World Building with a Generative AIHai Dang, Frederik Brudy, George W. Fitzmaurice, Fraser AndersonUIST 2023 · 43 citations
- Interaction-Augmented Instruction: Modeling the Synergy of Prompts and Interactions in Human-GenAI CollaborationLeixian Shen, Yifang Wang, Huamin Qu, Xing Xie et al.CHI 2026 · 3 citations
- Selectively Extracting and Injecting Visual Attributes into Text-to-Image ModelsSeunghwan Choi, Jooyeol Yun, Youngdo Lee, Jaegul ChooCVPR 2026
