SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
Sashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao, Ruofan Hu, Ziang Zhang, Xiaoda Yang, Zhibin Wang, Jun Song, Cheng Yu, Bo Zheng, Zhou Zhao
Abstract
Recent advances in text-to-image (T2I) generation via reinforcement learning (RL) have benefited from reward models that assess semantic alignment and visual quality. However, most existing reward models pay limited attention to fine-grained spatial relationships, often producing images that appear plausible overall yet contain inaccuracies in object positioning. In this work, we present SpatialReward, a verifiable reward model explicitly designed to evaluate spatial layouts in generated images. SpatialReward adopts a multi-stage pipeline: a Prompt Decomposer extracts entities, attributes, and spatial metadata from free-form prompts; expert detectors provide accurate visual grounding of object positions and attributes; and a vision-language model applies chain-of-thought reasoning over grounded observations to assess complex spatial relations that are challenging for rule-based methods. To more comprehensively evaluate spatial relationships in generated images, we introduce SpatRelBench, a benchmark covering object attributes, orientation, inter-object relations, and rendered text placement. Experiments on Stable Diffusion and FLUX show that incorporating SpatialReward into RL training consistently improves spatial consistency and overall generation quality, with results aligned more closely to human judgments. These findings indicate that verifiable reward models hold considerable potential for enabling more accurate and controllable optimization in text-to-image generation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac714468-3552-48dc-99a9-68e32471af9eCited by top-tier papers5
- GIFT: Global Irreplaceability Frame Targeting for Efficient Video UnderstandingJunpeng Ma, Sashuai Zhou, Guanghao Li, Xin Gao et al.CVPR 2026 · 7 citations
- GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View SynthesisXuqin Wang, Tao Wu, Yanfeng Zhang, Lu Liu et al.CVPR 2026 · 4 citations
- Hierarchical Codec Diffusion for Video-to-Speech GenerationJiaxin Ye, Gaoxiang Cong, Chenhui Wang, Xin-Cheng Wen et al.CVPR 2026 · 3 citations
- FontCrafter: High-Fidelity Element-Driven Artistic Font Creation with Visual In-Context GenerationWuyang Luo, Chengkaitan Chengkaitan to Chengkai Tan, Chang Ge, Binye Hong et al.CVPR 2026 · 2 citations
- Unified Thinker: A General Reasoning Core for Image GenerationSashuai Zhou, Qiang Zhou, Jijin Hu, Hanqing Yang et al.ACL 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Enhancing Spatial Understanding in Image Generation via Reward ModelingZhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu et al.CVPR 2026 · 2 citations
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image ModelsZengbin Wang, Xuecai Hu, Yong Wang, Feng Xiong et al.ICLR 2026 · 13 citations
- GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement LearningChengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang et al.ICLR 2026 · 43 citations
- SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial ReasoningYancheng Long, Yankai Yang, Hongyang Wei, Wei Chen et al.ICML 2026 · 6 citations
- SpotActor: Training-Free Layout-Controlled Consistent Image GenerationJiahao Wang, Caixia Yan, Weizhan Zhang, Haonan Lin et al.AAAI 2025 · 13 citations
