Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport
Lvmin Zhang, Anyi Rao, Maneesh Agrawala
摘要
Diffusion-based image generators are becoming unique methods for illumination harmonization and editing. The current bottleneck in scaling up the training of diffusion-based illumination editing models is mainly in the difficulty of preserving the underlying image details and maintaining intrinsic properties, such as albedos, unchanged. Without appropriate constraints, directly training the latest large image models with complex, varied, or in-the-wild data is likely to produce a structure-guided random image generator, rather than achieving the intended goal of precise illumination manipulation. We propose Imposing Consistent Light (IC-Light) transport during training, rooted in the physical principle that the linear blending of an object's appearances under different illumination conditions is consistent with its appearance under mixed illumination. This consistency allows for stable and scalable illumination learning, uniform handling of various data sources, and facilitates a physically grounded model behavior that modifies only the illumination of images while keeping other intrinsic properties unchanged. Based on this method, we can scale up the training of diffusion-based illumination editing models to large data quantities (>10 million), across all available data types (real light stages, rendered samples, in-the-wild synthetic augmentations, etc.), and using strong backbones (SDXL, Flux, etc.). We also demonstrate that this approach reduces uncertainties and mitigates artifacts such as mismatched materials or altered albedos.
Editing the illumination in images is a fundamental task in deep learning and image editing. Classic computer graphics methods often model the appearance of images using physical illumination models. More recently, large diffusion-based image generators have introduced unique applications and flexible paradigms in this area, handling a wider range of "in-the-wild" lighting effects beyond simply changing the distribution of light sources, e.g., generating backlighting or rim light, adding special effects like glow, glare, or the Tyndall effect, simulating shadows cast through tree shade or venetian blinds, and even manipulating human-drawn, composed, artistic, or non-photorealistic lighting conditions. These applications also provide tools for artists and designers to modify the foreground or background (e.g., product images, commercial posters, etc.) while maintaining harmonious illumination. These illumination editing applications with generative image models hold unique industrial value for visual content creation and manipulation.
Diffusion-based illumination editing methods also present new opportunities and considerations for scaling up training and utilizing stronger backbones. Yet, training an illumination editing model at larger scales and with more diversity is more challenging than it seems. The first challenge lies in maintaining the desired model behavior to ensure proper illumination manipulation rather than deviating into unintended random behaviors. As the dataset size and diversity increase, the mapping and distribution of the learning objective can become ambiguous and uncertain. Without appropriate constraints, the training may produce a structure-guided random image generator, resulting in outputs that do not align with the desired illumination editing requirements. This deviation occurs when the model fails to learn a mapping corresponding to illumination modification, instead introducing
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper77
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng 等NeurIPS 2025 · 被引用 45 次
- UniRelight: Learning Joint Decomposition and Synthesis for Video RelightingKai He, Ruofan Liang, Jacob Munkberg, Jon Hasselgren 等NeurIPS 2025 · 被引用 42 次
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningHao Tan, Jun Lan, Zichang Tan, Senyuan Shi 等ICLR 2026 · 被引用 26 次
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face ReconstructionSimon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito 等ICLR 2026 · 被引用 24 次
- GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDRChristophe Bolduc, Yannick Hold-Geoffroy, Jean-François LalondeICCV 2025 · 被引用 14 次
它引用的顶会 Paper57
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- TransLight: Image-Guided Customized Lighting Control with Generative DecouplingZongming Li, Lianghui Zhu, Haocheng Shen, Longjin Ran 等ICML 2026 · 被引用 2 次
- I 2HDiffuser: Image Illumination Harmonization Meets the Diffusion ModelZhongyun Bao, Gang Fu, Jianchi Sun, Jing Zhou 等ACM MM 2025 · 被引用 3 次
- Deep Image-based Illumination HarmonizationZhongyun Bao, Chengjiang Long, Gang Fu, Daquan Liu 等CVPR 2022 · 被引用 24 次
- Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion TransformersShuo Zhang, Wenzhuo Wu, Huayu Zhang, Jiarong Cheng 等ICLR 2026 · 被引用 1 次
- Relightful Harmonization: Lighting-Aware Portrait Background ReplacementMengwei Ren, Wei Xiong, Jae Shin Yoon, Zhixin Shu 等CVPR 2024 · 被引用 17 次
