DreamBooth++: Boosting Subject-Driven Generation via Region-Level References Packing
Zhongyi Fan, Zixin Yin, Gang Li, Yibing Zhan, Heliang Zheng
摘要
DreamBooth has demonstrated significant potential in subject-driven text-to-image generation, especially in scenarios requiring precise preservation of a subject's appearance. However, it still suffers from inefficiency and requires extensive iterative training to customize concepts using a small set of reference images. To address these issues, we introduce DreamBooth++, a region-level training strategy designed to significantly improve the efficiency and effectiveness of learning specific subjects. In particular, our approach employs a region-level data re-formulation technique that packs a set of reference images into a single sample, significantly reducing computational costs. Moreover, we adapt convolution and self-attention layers to ensure their processings are restricted within individual regions. Thus their operational scope (i.e., receptive field) can be preserved within a single subject, avoiding generating multiple sub-images within a single image. Last but not least, we design a text-guided prior regularization between our model and the pretrained one to preserve the original semantic generation ability. Comprehensive experiments demonstrate that our training strategy not only accelerates the subject-learning process but also significantly boosts fidelity to both subject and prompts in subject-driven generation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Personalized Visual Content Generation in Conversational SystemsXianquan Wang, Zhaocheng Du, Huibo Xu, Shukang Yin 等NeurIPS 2025 · 被引用 4 次
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture InfillingShuhong Zheng, Ashkan Mirzaei, Igor GilitschenskiNeurIPS 2025 · 被引用 2 次
- TARA: Token-Aware LoRA for Composable Personalization in Diffusion ModelsYuqi Peng, Lingtao Zheng, Yufeng Yang, Yi Huang 等AAAI 2026 · 被引用 2 次
- Fine-Tuning Visual Autoregressive Models for Subject-Driven GenerationJiwoo Chung, Sangeek Hyun, Hyunjun Kim, Eunseo Koh 等ICCV 2025 · 被引用 1 次
- HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait SynthesisMengtian Li, Jinshu Chen, Wanquan Feng, Bingchuan Li 等CVPR 2025
相关 Paper
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu 等NeurIPS 2024 · 被引用 18 次
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten 等SIGGRAPH 2024 · 被引用 57 次
- Storybooth: Training-Free Multi-Subject Consistency for Improved Visual StorytellingJaskirat Singh, Junshen K. Chen, Jonas Kohler, Michael F. CohenICLR 2025
- DreamBooth3D: Subject-Driven Text-to-3D GenerationAmit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer 等ICCV 2023 · 被引用 280 次
- MotionBooth: Motion-Aware Customized Text-to-Video GenerationJianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang 等NeurIPS 2024 · 被引用 114 次
