Subject-Diffusion: Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuning
Jian Ma, Junhao Liang, Chen Chen, Haonan Lu
摘要
Recent progress in personalized image generation using diffusion models has been significant. However, development in the area of open-domain and test-time fine-tuning-free personalized image generation is proceeding rather slowly. In this paper, we propose Subject-Diffusion, a novel open-domain personalized image generation model that, in addition to not requiring test-time fine-tuning, also only requires a single reference image to support personalized generation of single- or two-subjects in any domain. Firstly, we construct an automatic data labeling tool and use the LAION-Aesthetics dataset to construct a large-scale dataset consisting of 76M images and their corresponding subject detection bounding boxes, segmentation masks, and text descriptions. Secondly, we design a new unified framework that combines text and image semantics by incorporating coarse location and fine-grained reference image control to maximize subject fidelity and generalization. Furthermore, we also adopt an attention control mechanism to support two-subject generation. Extensive qualitative and quantitative results demonstrate that our method have certain advantages over other frameworks in single, multiple, and human-customized image generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper87
- Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model EvaluationShih-Ying Yeh, Yu-Guan Hsieh, Zhidong Gao, Bernard B. W. Yang 等ICLR 2024 · 被引用 133 次
- CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition AbilitiesTao Wu, Yong Zhang, Xintao Wang, Xianpan Zhou 等AAAI 2025 · 被引用 62 次
- Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image GenerationQihan Huang, Siming Fu, Jinlong Liu, Hao Jiang 等AAAI 2025 · 被引用 43 次
- DiLightNet: Fine-grained Lighting Control for Diffusion-based Image GenerationChong Zeng, Yue Dong, Pieter Peers, Youkang Kong 等SIGGRAPH 2024 · 被引用 41 次
- OminiControl: Minimal and Universal Control for Diffusion TransformerZhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue 等ICCV 2025 · 被引用 34 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Identity Decoupling for Multi-Subject Personalization of Text-to-Image ModelsSangwon Jang, Jaehyeong Jo, Kimin Lee, Sung Ju HwangNeurIPS 2024 · 被引用 42 次
- Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image PersonalizationHenglei Lv, Jiayu Xiao, Liang LiACM MM 2024 · 被引用 6 次
- DynASyn: Multi-Subject Personalization Enabling Dynamic Action SynthesisYongjin Choi, Chanhun Park, Seung Jun BaekAAAI 2025 · 被引用 3 次
- ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven GenerationDaniel Winter, Asaf Shul, Matan Cohen, Dana Berman 等ICCV 2025 · 被引用 2 次
- MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout GuidanceXierui Wang, Siming Fu, Qihan Huang, Wanggui He 等ICLR 2025
