High-fidelity Person-centric Subject-to-Image Synthesis
Yibin Wang, Weizhong Zhang, Jianwei Zheng, Cheng Jin
摘要
Current subject-driven image generation methods en-counter significant challenges in person-centric image generation. The reason is that they learn the semantic scene and person generation by fine-tuning a common pre-trained diffusion, which involves an irreconcilable training imbalance. Precisely, to generate realistic persons, they need to sufficiently tune the pre-trained model, which inevitably causes the model to forget the rich semantic scene prior and makes scene generation over-fit to the training data. Moreover, even with sufficient fine-tuning, these methods can still not generate high-fidelity persons since joint learning of the scene and person generation also lead to quality compromise. In this paper, we propose Face-diffuser, an effective collaborative generation pipeline to eliminate the above training imbal-ance and quality compromise. Specifically, we first develop two specialized pre-trained diffusion models, i.e., Text-driven Diffusion Model (TDM) and Subject-augmented Diffusion Model (SDM), for scene and person generation, respectively. The sampling process is divided into three sequential stages, i.e., semantic scene construction, subject-scene fusion, and subject enhancement. The first and last stages are performed by TDM and SDM respectively. The subject-scene fusion stage, that is the collaboration achieved through a novel and highly effective mechanism, Saliency-adaptive Noise Fusion (SNF). Specifically, it is based on our key observation that there exists a robust link between classifier-free guidance responses and the saliency of generated images. In each time step, SNF leverages the unique strengths of each model and allows for the spatial blending of predicted noises from both models automatically in a saliency-aware manner, all of which can be seamlessly integrated into the DDIM sampling process. Extensive experiments confirm the impressive effectiveness and robustness of the Face-diffuser in gener-ating high-fidelity person images depicting multiple unseen persons with varying contexts. Code is available at https://github.com/CodeGoat24/Face-diffuser.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- WithAnyone: Toward Controllable and ID Consistent Image GenerationHengyuan Xu, Wei Cheng, Peng Xing, Yixiao Fang 等ICLR 2026 · 被引用 12 次
- PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention SteeringYibin Wang, Weizhong Zhang, Jianwei Zheng, Cheng JinACM MM 2024 · 被引用 9 次
- Importance-Based Token Merging for Efficient Image and Video GenerationHaoyu Wu, Jingyi Xu, Hieu Le, Dimitris SamarasICCV 2025 · 被引用 3 次
- DynamicID: Zero-Shot Multi-ID Image Personalization With Flexible Facial EditabilityXirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang 等ICCV 2025 · 被引用 3 次
- Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation KnowledgeYiyang Cai, Zhengkai Jiang, Yulong Liu, Chunyang Jiang 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationYuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai 等ICCV 2023 · 被引用 469 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
相关 Paper
- MagicFusion: Boosting Text-to-Image Generation Performance by Fusing Diffusion ModelsJing Zhao, Heliang Zheng, Chaoyue Wang, Long Lan 等ICCV 2023 · 被引用 25 次
- Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image SynthesisYanzuo Lu, Manlin Zhang, Andy J. Ma, Xiaohua Xie 等CVPR 2024 · 被引用 26 次
- Face2Diffusion for Fast and Editable Face PersonalizationKaede Shiohara, Toshihiko YamasakiCVPR 2024 · 被引用 13 次
- Teleportraits: Training-Free People Insertion Into Any SceneJialu Gao, K. J. Joseph, Fernando De la TorreICCV 2025
- Steering Guidance for Personalized Text-to-Image Diffusion ModelsSunghyun Park, Seokeon Choi, Hyoungwoo Park, Sungrack YunICCV 2025 · 被引用 2 次
