Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modeling
Jianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li, Zhenyu Tang, Wangbo Yu, Yian Zhao, Xiaomin Li, Li Yuan, Yonghong Tian
摘要
We present a novel approach to bridging the gap between style transfer guided by reference styles and faithful content preservation in text-to-image generation. First, we introduce StyleReward-Dataset, an expert-annotated benchmark with over 300k adversarial image pairs spanning diverse real-world and virtual styles. Each example is constructed to contrast a faithful result with targeted counterexamples, enabling preference learning over style consistency, content preservation, and perceptual quality. Leveraging StyleReward-Dataset, we propose Style-Score, an end-to-end multimodal reward model that integrates visual and semantic understanding to provide reliable, human-aligned evaluations of generated images. Additionally, based on StyleReward-Dataset, we develop a two-stage training framework for style transfer, consisting of supervised fine-tuning on StyleReward-Dataset followed by reinforcement learning with Group Relative Preference Optimization that converts relative preferences into stable policy updates. Extensive experiments on both public benchmark and proposed benchmark show that our proposed method achieves substantial improvements in style fidelity and content preservation over strong baselines, with human studies further validating its superior visual realism and alignment with human preference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- StyleDoctor: Towards Specialist Reward Model for Style-centric Generation TasksXilin He, Xiaole Xian, Xiangyu Yue, Muhammad Haris KhanCVPR 2026
- Science-T2I: Addressing Scientific Illusions in Image SynthesisJialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu 等CVPR 2025
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- Enhancing Spatial Understanding in Image Generation via Reward ModelingZhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu 等CVPR 2026 · 被引用 2 次
- Curriculum Direct Preference Optimization for Diffusion and Consistency ModelsFlorinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, Nicu Sebe 等CVPR 2025
