Enhancing Spatial Understanding in Image Generation via Reward Modeling
Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu, Xiaojie Li, Rui Wang, Yunpeng Chen, Daquan Zhou
Abstract
Recent progress in text-to-image generation has greatly advanced visual fidelity and creativity, but it has also imposed higher demands on prompt complexity-particularly in encoding intricate spatial relationships. In such cases, achieving satisfactory results often requires multiple sampling attempts. To address this challenge, we introduce a novel method that strengthens the spatial understanding of current image generation models. We first construct the SpatialReward-Dataset with over 80k preference pairs. Building on this dataset, we build SpatialScore, a reward model designed to evaluate the accuracy of spatial relationships in text-to-image generation, achieving performance that even surpasses leading proprietary models on spatial evaluation. We further demonstrate that this reward model effectively enables online reinforcement learning for the complex spatial generation. Extensive experiments across multiple benchmarks show that our specialized reward model yields significant and consistent gains in spatial understanding for image generation. The visual demo is available at the project page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial ReasoningYancheng Long, Yankai Yang, Hongyang Wei, Wei Chen et al.ICML 2026 · 6 citations
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao et al.CVPR 2026 · 7 citations
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward ModelingJianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li et al.CVPR 2026
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image ModelsZengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong et al.ICLR 2026 · 13 citations
- GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement LearningChengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang et al.ICLR 2026 · 43 citations
