Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models
Gihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee, Jong Chul Ye, Fabian Caba Heilbron
摘要
While there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this work, we introduce Concept Weaver, a method for composing customized text-to-image diffusion models at inference time. Specifically, the method breaks the process into two steps: creating a template image aligned with the semantics of input prompts, and then personalizing the template using a concept fusion strategy. The fusion strategy incorporates the appearance of the target concepts into the template image while retaining its structural details. The results indicate that our method can generate multiple custom concepts with higher identity fidelity compared to alternative approaches. Furthermore, the method is shown to seamlessly handle more than two concepts and closely follow the semantic meaning of the input prompt without blending appearances across different subjects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- An Efficient and Harmonized Framework for Balanced Cross-Domain Feature IntegrationShaoxu Li, Ye PanAAAI 2026 · 被引用 15 次
- Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image SynthesisZebin Yao, Fangxiang Feng, Ruifan Li, Xiaojie WangAAAI 2025 · 被引用 3 次
- Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view DiffusionZeren Xiong, Zikun Chen, Zedong Zhang, Xiang Li 等ACM MM 2025 · 被引用 2 次
- UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image PersonalizationJunjie He, Yifeng Geng, Liefeng BoICCV 2025 · 被引用 2 次
- Continual Personalization for Diffusion ModelsYu-Chien Liao, Jr-Jen Chen, Chi-Pin Huang, Ci-Siang Lin 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored PromptsFeng Liang, Haoyu Ma, Zecheng He, Tingbo Hou 等CVPR 2025
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 被引用 13 次
- DynASyn: Multi-Subject Personalization Enabling Dynamic Action SynthesisYongjin Choi, Chanhun Park, Seung Jun BaekAAAI 2025 · 被引用 3 次
- TokenVerse: Versatile Multi-concept Personalization in Token Modulation SpaceDaniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov 等SIGGRAPH 2025 · 被引用 12 次
- MC^2: Multi-concept Guidance for Customized Multi-concept GenerationJiaxiu Jiang, Yabo Zhang, Kailai Feng, Xiaohe Wu 等CVPR 2025
