UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
Pengwei Liu, Hangjie Yuan, Bo Dong, Jiazheng Xing, Jinwang Wang, Rui Zhao, Weihua Chen, Fan Wang
Abstract
Relighting is a crucial task with both practical demand and artistic value, and recent diffusion models have shown strong potential by enabling rich and controllable lighting effects. However, as they are typically optimized in semantic latent space, where proximity does not guarantee physical correctness in visual space, they often produce unrealistic results, such as overexposed highlights, misaligned shadows, and incorrect occlusions. We address this with UniLumos, a unified relighting framework for both images and videos that brings RGB-space geometry feedback into a flow matching backbone. By supervising the model with depth and normal maps extracted from its outputs, we explicitly align lighting effects with the scene structure, enhancing physical plausibility. Nevertheless, this feedback requires high-quality outputs for supervision in visual space, making standard multi-step denoising computationally expensive. To mitigate this, we employ path consistency learning, allowing supervision to remain effective even under few-step training regimes. To enable fine-grained relighting control and supervision, we design a structured six-dimensional annotation protocol capturing core illumination attributes. Building upon this, we propose LumosBench, a disentangled attribute-level benchmark that evaluates lighting controllability via large vision-language models, enabling automatic and interpretable assessment of relighting precision across individual dimensions. Extensive experiments demonstrate that UniLumos achieves state-of-the-art relighting quality with significantly improved physical consistency, while delivering a 20x speedup for both image and video relighting. Code is available at https://github.com/alibaba-damo-academy/Lumos-Custom.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 976cbfb3-4446-4416-8daa-a255a1d1680bCited by top-tier papers2
- Learning Latent Proxies for Controllable Single-Image RelightingHaoze Zheng, Zihao Wang, Xianfeng Wu, Yajing Bai et al.CVPR 2026 · 1 citation
- TokenLight: Precise Lighting Control in Images using Attribute TokensSumit Chaturvedi, Yannick Hold-Geoffroy, Mengwei Ren, Jingyuan Liu et al.CVPR 2026 · 1 citation
Builds on34
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
Related papers
- UniRelight: Learning Joint Decomposition and Synthesis for Video RelightingKai He, Ruofan Liang, Jacob Munkberg, Jon Hasselgren et al.NeurIPS 2025 · 42 citations
- Relit-LiVE: Relight Video by Jointly Learning Environment VideoWeiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen et al.SIGGRAPH 2026
- Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and HarmonizationJunying Wang, Jingyuan Liu, Xin Sun, Krishna Kumar Singh et al.CVPR 2025
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu et al.ICLR 2026 · 21 citations
- IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video GenerationYuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai, Ronald Clark et al.NeurIPS 2025 · 6 citations
