Lune

CVPR2026顶会

Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modeling

Jianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li, Zhenyu Tang, Wangbo Yu, Yian Zhao, Xiaomin Li, Li Yuan, Yonghong Tian

出版方
2026年份

摘要

We present a novel approach to bridging the gap between style transfer guided by reference styles and faithful content preservation in text-to-image generation. First, we introduce StyleReward-Dataset, an expert-annotated benchmark with over 300k adversarial image pairs spanning diverse real-world and virtual styles. Each example is constructed to contrast a faithful result with targeted counterexamples, enabling preference learning over style consistency, content preservation, and perceptual quality. Leveraging StyleReward-Dataset, we propose Style-Score, an end-to-end multimodal reward model that integrates visual and semantic understanding to provide reliable, human-aligned evaluations of generated images. Additionally, based on StyleReward-Dataset, we develop a two-stage training framework for style transfer, consisting of supervised fine-tuning on StyleReward-Dataset followed by reinforcement learning with Group Relative Preference Optimization that converts relative preferences into stable policy updates. Extensive experiments on both public benchmark and proposed benchmark show that our proposed method achieves substantial improvements in style fidelity and content preservation over strong baselines, with human studies further validating its superior visual realism and alignment with human preference.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper35

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖