Enhancing Spatial Understanding in Image Generation via Reward Modeling
Zhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu, Xiaojie Li, Rui Wang, Yunpeng Chen, Daquan Zhou
摘要
Recent progress in text-to-image generation has greatly advanced visual fidelity and creativity, but it has also imposed higher demands on prompt complexity-particularly in encoding intricate spatial relationships. In such cases, achieving satisfactory results often requires multiple sampling attempts. To address this challenge, we introduce a novel method that strengthens the spatial understanding of current image generation models. We first construct the SpatialReward-Dataset with over 80k preference pairs. Building on this dataset, we build SpatialScore, a reward model designed to evaluate the accuracy of spatial relationships in text-to-image generation, achieving performance that even surpasses leading proprietary models on spatial evaluation. We further demonstrate that this reward model effectively enables online reinforcement learning for the complex spatial generation. Extensive experiments across multiple benchmarks show that our specialized reward model yields significant and consistent gains in spatial understanding for image generation. The visual demo is available at the project page.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial ReasoningYancheng Long, Yankai Yang, Hongyang Wei, Wei Chen 等ICML 2026 · 被引用 6 次
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao 等CVPR 2026 · 被引用 7 次
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward ModelingJianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li 等CVPR 2026
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image ModelsZengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong 等ICLR 2026 · 被引用 13 次
- GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement LearningChengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang 等ICLR 2026 · 被引用 43 次
