Color Alignment in Diffusion
Ka-Chun Shum, Binh-Son Hua, Duc Thanh Nguyen, Sai-Kit Yeung
摘要
Diffusion models have shown great promise in synthesizing visually appealing images. However, it remains challenging to condition the synthesis at a fine-grained level, for instance, synthesizing image pixels following some generic color pattern. Existing image synthesis methods often produce contents that fall outside the desired pixel conditions. To address this, we introduce a novel color alignment algorithm that confines the generative process in diffusion models within a given color pattern. Specifically, we project diffusion terms, either imagery samples or latent representations, into a conditional color space to align with the input color distribution. This strategy simplifies the prediction in diffusion models within a color manifold while still allowing plausible structures in generated contents, thus enabling the generation of diverse contents that comply with the target color pattern. Experimental results demonstrate our stateof-the-art performance in conditioning and controlling of color pixels, while maintaining on-par generation quality and diversity in comparison with regular diffusion models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Color Conditional Generation with Sliced Wasserstein GuidanceAlexander Lobashev, Maria A. Larchenko, Dmitry GuskovNeurIPS 2025 · 被引用 11 次
- PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic TrajectoriesGemma Canet Tarrés, Manel Baradad, Francesc Moreno-Noguer, Yumeng LiCVPR 2026 · 被引用 1 次
- The Latent Color Subspace: Emergent Order in High-Dimensional ChaosMateusz Pach, Jessica Bader, Quentin Bouniot, Serge Belongie 等ICML 2026
- Colorful-Noise: Training-Free Low-Frequency Noise Manipulation for Color-Based Conditional Image GenerationNadav Z. Cohen, Ofir Abramovich, Ariel ShamirSIGGRAPH 2026
- Gradient Descent in the ALPS: Abstracted Low-Poly Stylization and FabricationRuben Wiersma, Alexandre Binninger, Peizhuo Li, Tanguy Magne 等SIGGRAPH 2026
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image GenerationZijing Hu, Yunze Tong, Fengda Zhang, Junkun Yuan 等ICLR 2026 · 被引用 3 次
- Exploring Palette based Color Guidance in Diffusion ModelsQianru Qiu, Jiafeng Mao, Xueting WangACM MM 2025 · 被引用 4 次
- Real-World Image Variation by Aligning Diffusion Inversion ChainYuechen Zhang, Jinbo Xing, Eric Lo, Jiaya JiaNeurIPS 2023 · 被引用 56 次
- Style Aligned Image Generation via Shared AttentionAmir Hertz, Andrey Voynov, Shlomi Fruchter, Daniel Cohen-OrCVPR 2024
- Inference-Time Alignment of Diffusion Models with Direct Noise OptimizationZhiwei Tang, Jiangweizhi Peng, Jiasheng Tang, Mingyi Hong 等ICML 2025
