StyleDrop: Text-to-Image Synthesis of Any Style
Kihyuk Sohn, Lu Jiang, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan Essa, Michael Rubinstein, Yuan Hao, Glenn Entis
Abstract
Visualization of StyleDrop outputs generated by personalized text-to-image models for 18 different styles. Each model is tuned on a single style reference image, which is shown in the white insert box of each image. The per-style text descriptor is appended to the content text prompt: "A fluffy baby sloth with a knitted hat trying to figure out a laptop, close up". Generated images capture many nuances such as colors, shading, textures and 3D appearance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Identity Decoupling for Multi-Subject Personalization of Text-to-Image ModelsSangwon Jang, Jaehyeong Jo, Kimin Lee, Sung Ju HwangNeurIPS 2024 · 42 citations
- Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story GenerationAo Ma, Jiasong Feng, Ke Cao, Jing Wang et al.ICCV 2025 · 13 citations
- Color Conditional Generation with Sliced Wasserstein GuidanceAlexander Lobashev, Maria A. Larchenko, Dmitry GuskovNeurIPS 2025 · 11 citations
- CSD-VAR: Content-Style Decomposition in Visual Autoregressive ModelsQuang-Binh Nguyen, Minh Luu, Quang Nguyen, Anh Tran et al.ICCV 2025 · 7 citations
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- TextCraftor: Your Text Encoder can be Image Quality ControllerYanyu Li, Xian Liu, Anil Kag, Ju Hu et al.CVPR 2024
- StyleStudio: Text-Driven Style Transfer with Selective Control of Style ElementsMingkun Lei, Xue Song, Beier Zhu, Hao Wang et al.CVPR 2025
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- Magic Insert: Style-Aware Drag-And-DropNataniel Ruiz, Yuanzhen Li, Neal Wadhwa, Yael Pritch et al.ICCV 2025
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik et al.ICLR 2023 · 464 citations
