Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling
Guiyu Zhang, Huan-ang Gao, Zijian Jiang, Hao Zhao, Zhedong Zheng
Abstract
In this paper, we focus on the task of conditional image generation, where an image is synthesized according to user instructions. The critical challenge underpinning this task is ensuring both the fidelity of the generated images and their semantic alignment with the provided conditions. To tackle this issue, previous studies have employed supervised perceptual losses derived from pre-trained models, i.e., reward models, to enforce alignment between the condition and the generated result. However, we observe one inherent shortcoming: considering the diversity of synthesized images, the reward model usually provides inaccurate feedback when encountering newly generated data, which can undermine the training process. To address this limitation, we propose an uncertainty-aware reward modeling, called Ctrl-U, including uncertainty estimation and uncertaintyaware regularization, designed to reduce the adverse effects of imprecise feedback from the reward model. Given the inherent cognitive uncertainty within reward models, even images generated under identical conditions often result in a relatively large discrepancy in reward loss. Inspired by the observation, we explicitly leverage such prediction variance as an uncertainty indicator. Based on the uncertainty estimation, we regularize the model training by adaptively rectifying the reward. In particular, rewards with lower uncertainty receive higher loss weights, while those with higher uncertainty are given reduced weights to allow for larger variability. The proposed uncertainty regularization facilitates reward fine-tuning through consistency construction. Extensive experiments validate the effectiveness of our methodology in improving the controllability and generation quality, as well as its scalability across diverse conditional scenarios, including segmentation mask, edge, and depth conditions. Codes are publicly available at https://grenoble-zhang.github.io/Ctrl-U .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ba23a9c-b233-4262-85d7-3bc0b365ea19Cited by top-tier papers6
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual ContextsYuchen Zhang, Yaxiong Wang, Yujiao Wu, Lianwei Wu et al.CVPR 2026 · 8 citations
- Generating Attribution Reports for Manipulated Facial Images: A Dataset and BaselineJingchun Lian, Lingyu Liu, Yaxiong Wang, Yujiao Wu et al.ACL 2026 · 7 citations
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li et al.CVPR 2026 · 3 citations
- Discovering Latent Graphs with GFlowNets for Diverse Conditional Image GenerationBailey Trang Nguyen, Parham Saremi, Alan Q. Wang, Fangrui Huang et al.NeurIPS 2025 · 2 citations
- PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction ModelMingju Gao, Yike Pan, Huan-ang Gao, Zongzheng Zhang et al.CVPR 2025
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Learn to Guide Your Diffusion ModelAlexandre Galashov, Ashwini Pokle, Arnaud Doucet, Arthur Gretton et al.ICLR 2026 · 12 citations
- SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive ModelDongli Xu, Aleksei Tiulpin, Matthew B. BlaschkoICLR 2026 · 2 citations
- Efficient Tail-Aware Generative Optimization via Flow Model Fine-TuningZifan Wang, Riccardo De Santi, Xiaoyu Mo, Michael Zavlanos et al.ICML 2026 · 4 citations
- Improving Instruction Following in Language Models through Proxy-Based Uncertainty EstimationJoonHo Lee, Jae Oh Woo, Juree Seok, Parisa Hassanzadeh et al.ICML 2024 · 4 citations
- Censored Sampling of Diffusion Models Using 3 Minutes of Human FeedbackTaeho Yoon, Kibeom Myoung, Keon Lee, Jaewoong Cho et al.NeurIPS 2023 · 13 citations
