Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport
Lvmin Zhang, Anyi Rao, Maneesh Agrawala
Abstract
Diffusion-based image generators are becoming unique methods for illumination harmonization and editing. The current bottleneck in scaling up the training of diffusion-based illumination editing models is mainly in the difficulty of preserving the underlying image details and maintaining intrinsic properties, such as albedos, unchanged. Without appropriate constraints, directly training the latest large image models with complex, varied, or in-the-wild data is likely to produce a structure-guided random image generator, rather than achieving the intended goal of precise illumination manipulation. We propose Imposing Consistent Light (IC-Light) transport during training, rooted in the physical principle that the linear blending of an object's appearances under different illumination conditions is consistent with its appearance under mixed illumination. This consistency allows for stable and scalable illumination learning, uniform handling of various data sources, and facilitates a physically grounded model behavior that modifies only the illumination of images while keeping other intrinsic properties unchanged. Based on this method, we can scale up the training of diffusion-based illumination editing models to large data quantities (>10 million), across all available data types (real light stages, rendered samples, in-the-wild synthetic augmentations, etc.), and using strong backbones (SDXL, Flux, etc.). We also demonstrate that this approach reduces uncertainties and mitigates artifacts such as mismatched materials or altered albedos.
Editing the illumination in images is a fundamental task in deep learning and image editing. Classic computer graphics methods often model the appearance of images using physical illumination models. More recently, large diffusion-based image generators have introduced unique applications and flexible paradigms in this area, handling a wider range of "in-the-wild" lighting effects beyond simply changing the distribution of light sources, e.g., generating backlighting or rim light, adding special effects like glow, glare, or the Tyndall effect, simulating shadows cast through tree shade or venetian blinds, and even manipulating human-drawn, composed, artistic, or non-photorealistic lighting conditions. These applications also provide tools for artists and designers to modify the foreground or background (e.g., product images, commercial posters, etc.) while maintaining harmonious illumination. These illumination editing applications with generative image models hold unique industrial value for visual content creation and manipulation.
Diffusion-based illumination editing methods also present new opportunities and considerations for scaling up training and utilizing stronger backbones. Yet, training an illumination editing model at larger scales and with more diversity is more challenging than it seems. The first challenge lies in maintaining the desired model behavior to ensure proper illumination manipulation rather than deviating into unintended random behaviors. As the dataset size and diversity increase, the mapping and distribution of the learning objective can become ambiguous and uncertain. Without appropriate constraints, the training may produce a structure-guided random image generator, resulting in outputs that do not align with the desired illumination editing requirements. This deviation occurs when the model fails to learn a mapping corresponding to illumination modification, instead introducing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers77
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- UniRelight: Learning Joint Decomposition and Synthesis for Video RelightingKai He, Ruofan Liang, Jacob Munkberg, Jon Hasselgren et al.NeurIPS 2025 · 42 citations
- Veritas: Generalizable Deepfake Detection via Pattern-Aware ReasoningHao Tan, Jun Lan, Zichang Tan, Senyuan Shi et al.ICLR 2026 · 26 citations
- Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face ReconstructionSimon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito et al.ICLR 2026 · 24 citations
- GaSLight: Gaussian Splats for Spatially-Varying Lighting in HDRChristophe Bolduc, Yannick Hold-Geoffroy, Jean-François LalondeICCV 2025 · 14 citations
Builds on57
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- TransLight: Image-Guided Customized Lighting Control with Generative DecouplingZongming Li, Lianghui Zhu, Haocheng Shen, Longjin Ran et al.ICML 2026 · 2 citations
- I 2HDiffuser: Image Illumination Harmonization Meets the Diffusion ModelZhongyun Bao, Gang Fu, Jianchi Sun, Jing Zhou et al.ACM MM 2025 · 3 citations
- Deep Image-based Illumination HarmonizationZhongyun Bao, Chengjiang Long, Gang Fu, Daquan Liu et al.CVPR 2022 · 24 citations
- Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion TransformersShuo Zhang, Wenzhuo Wu, Huayu Zhang, Jiarong Cheng et al.ICLR 2026 · 1 citation
- Relightful Harmonization: Lighting-Aware Portrait Background ReplacementMengwei Ren, Wei Xiong, Jae Shin Yoon, Zhixin Shu et al.CVPR 2024 · 17 citations
