Color Alignment in Diffusion
Ka-Chun Shum, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
Abstract
Diffusion models have shown great promise in synthesizing visually appealing images. However, it remains challenging to condition the synthesis at a fine-grained level, for instance, synthesizing image pixels following some generic color pattern. Existing image synthesis methods often produce contents that fall outside the desired pixel conditions. To address this, we introduce a novel color alignment algorithm that confines the generative process in diffusion models within a given color pattern. Specifically, we project diffusion terms, either imagery samples or latent representations, into a conditional color space to align with the input color distribution. This strategy simplifies the prediction in diffusion models within a color manifold while still allowing plausible structures in generated contents, thus enabling the generation of diverse contents that comply with the target color pattern. Experimental results demonstrate our stateof-the-art performance in conditioning and controlling of color pixels, while maintaining on-par generation quality and diversity in comparison with regular diffusion models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Color Conditional Generation with Sliced Wasserstein GuidanceAlexander Lobashev, Maria A. Larchenko, Dmitry GuskovNeurIPS 2025 · 11 citations
- PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic TrajectoriesGemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer, Yumeng LiCVPR 2026 · 1 citation
- The Latent Color Subspace: Emergent Order in High-Dimensional ChaosMateusz Pach, Jessica Bader, Quentin Bouniot, Serge Belongie et al.ICML 2026
- Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image GenerationNadav Z. Cohen, Ofir Abramovich, Ariel ShamirSIGGRAPH 2026
- Gradient Descent in the ALPS: Abstracted Low-Poly Stylization and FabricationRuben Wiersma, Alexandre Binninger, Peizhuo Li, Tanguy Magne et al.SIGGRAPH 2026
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan et al.ICLR 2026 · 3 citations
- Exploring Palette based Color Guidance in Diffusion ModelsQianru Qiu, Jiafeng Mao, Xueting WangACM MM 2025 · 4 citations
- Real-World Image Variation by Aligning Diffusion Inversion ChainYuechen Zhang, Jinbo Xing, Eric Lo, Jiaya JiaNeurIPS 2023 · 56 citations
- Style Aligned Image Generation via Shared AttentionAmir Hertz, Andrey Voynov, Shlomi Fruchter, Daniel Cohen-OrCVPR 2024
- Inference-Time Alignment of Diffusion Models with Direct Noise OptimizationZhiwei Tang, Jiangweizhi Peng, Jiasheng Tang, Mingyi Hong et al.ICML 2025
