Lune

ICLR2025Top-tier venue

Scaling In-the-Wild Training for Diffusion-based Illumination Harmonization and Editing by Imposing Consistent Light Transport

Lvmin Zhang, Anyi Rao, Maneesh Agrawala

2025Year
77Top-tier citations

Abstract

Diffusion-based image generators are becoming unique methods for illumination harmonization and editing. The current bottleneck in scaling up the training of diffusion-based illumination editing models is mainly in the difficulty of preserving the underlying image details and maintaining intrinsic properties, such as albedos, unchanged. Without appropriate constraints, directly training the latest large image models with complex, varied, or in-the-wild data is likely to produce a structure-guided random image generator, rather than achieving the intended goal of precise illumination manipulation. We propose Imposing Consistent Light (IC-Light) transport during training, rooted in the physical principle that the linear blending of an object's appearances under different illumination conditions is consistent with its appearance under mixed illumination. This consistency allows for stable and scalable illumination learning, uniform handling of various data sources, and facilitates a physically grounded model behavior that modifies only the illumination of images while keeping other intrinsic properties unchanged. Based on this method, we can scale up the training of diffusion-based illumination editing models to large data quantities (>10 million), across all available data types (real light stages, rendered samples, in-the-wild synthetic augmentations, etc.), and using strong backbones (SDXL, Flux, etc.). We also demonstrate that this approach reduces uncertainties and mitigates artifacts such as mismatched materials or altered albedos.

Editing the illumination in images is a fundamental task in deep learning and image editing. Classic computer graphics methods often model the appearance of images using physical illumination models. More recently, large diffusion-based image generators have introduced unique applications and flexible paradigms in this area, handling a wider range of "in-the-wild" lighting effects beyond simply changing the distribution of light sources, e.g., generating backlighting or rim light, adding special effects like glow, glare, or the Tyndall effect, simulating shadows cast through tree shade or venetian blinds, and even manipulating human-drawn, composed, artistic, or non-photorealistic lighting conditions. These applications also provide tools for artists and designers to modify the foreground or background (e.g., product images, commercial posters, etc.) while maintaining harmonious illumination. These illumination editing applications with generative image models hold unique industrial value for visual content creation and manipulation.

Diffusion-based illumination editing methods also present new opportunities and considerations for scaling up training and utilizing stronger backbones. Yet, training an illumination editing model at larger scales and with more diversity is more challenging than it seems. The first challenge lies in maintaining the desired model behavior to ensure proper illumination manipulation rather than deviating into unintended random behaviors. As the dataset size and diversity increase, the mapping and distribution of the learning objective can become ambiguous and uncertain. Without appropriate constraints, the training may produce a structure-guided random image generator, resulting in outputs that do not align with the desired illumination editing requirements. This deviation occurs when the model fails to learn a mapping corresponding to illumination modification, instead introducing

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers77

Ask how each one uses it

Builds on57

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines