EMControl: Adding Conditional Control to Text-to-Image Diffusion Models via Expectation-Maximization
He Wang, Longquan Dai, Jinhui Tang
摘要
Recent advances in diffusion models focus on efficiently handling conditional generative tasks without extra training. The process involves decomposing the result into two components: 1. unconditional sample, generated in the absence of conditions; 2. condition correction, adjusting unconditional sample to include the guidance image. This adjustment is quantified by the pixel-level measure, where the latent is decoded back into a pixel image, and the forward operator translates the noisy image into the guidance domain for comparison with the guidance image. To enhance the fidelity of condition correction, we propose a learnable latent forward operator, focusing on latent-space consistency with the expectation that this latent-space consistency approximates the pixel-level fidelity measure. The encoder translates the guidance image into the latent space, and a correctional operator is proposed to rectify model mismatching in the latent guidance model. The determination of the condition term and the correction estimation is akin to solving a blind inverse problem. Our EMControl employs the Expectation-Maximization (EM) algorithm to solve the blind inverse problem during the reverse sampling process. This technique ensures that samples, once consistent with the guidance, are accurately mapped back onto the noisy data manifold, adhering to the data's inherent distribution. The EMControl has proven its effectiveness by delivering superior performance in conditional diffusion generation tasks compared to previous approaches. Moreover, its application to multiple-condition scenarios underscores its versatility and robustness across a range of generative tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningJiawei Gu, Ziyue Qiao, Xinming Li, Zechao LiNeurIPS 2025 · 被引用 5 次
- Refining Norms: A Post-hoc Framework for OOD Detection in Graph Neural NetworksJiawei Gu, Ziyue Qiao, Zechao LiNeurIPS 2025 · 被引用 3 次
- Mitigating Query Selection Bias in Referring Video Object SegmentationDingwei Zhang, Dong Zhang, Jinhui TangACM MM 2025 · 被引用 1 次
- DISCO: DISCrete nOise for Conditional Control in Text-to-Image Diffusion ModelsLongquan Dai, Ming Wu, Dejiao Xue, He Wang 等NeurIPS 2025
- STEDiff: Revealing the Spatial and Temporal Redundancy of Backdoor Attacks in Text-to-Image Diffusion ModelsYu Pan, Jiahao Chen, Lin Wang, Bingrong Dai 等ICLR 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- MMD Guidance: Training-Free Distribution Adaptation for Diffusion Models via Maximum Mean Discrepancy GuidanceMatina Mahdizadeh Sani, Nima Jamali, Mohammad Jalali, Farzan FarniaICML 2026 · 被引用 4 次
- NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion ModelsLongquan Dai, He Wang, Jinhui TangCVPR 2025
- Conducting Conditional Diffusion by Estimating the Mean Vector of von Mises-Fisher DistributionLongquan Dai, He Wang, Xiaolu Wei, Shaomeng Wang 等ACM MM 2025
- Solving Inverse Problems with Latent Diffusion Models via Hard Data ConsistencyBowen Song, Soo Min Kwon, Zecheng Zhang, Xinyu Hu 等ICLR 2024 · 被引用 213 次
- DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image EditingZixiang Li, Haoyu Wang, Wei Wang, Chuangchuang Tan 等NeurIPS 2025 · 被引用 5 次
