Lune

CVPR2026Top-tier venue

Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward Modeling

Jianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li, Zhenyu Tang, Wangbo Yu, Yian Zhao, Xiaomin Li, Li Yuan, Yonghong Tian

2026Year

Abstract

We present a novel approach to bridging the gap between style transfer guided by reference styles and faithful content preservation in text-to-image generation. First, we introduce StyleReward-Dataset, an expert-annotated benchmark with over 300k adversarial image pairs spanning diverse real-world and virtual styles. Each example is constructed to contrast a faithful result with targeted counterexamples, enabling preference learning over style consistency, content preservation, and perceptual quality. Leveraging StyleReward-Dataset, we propose Style-Score, an end-to-end multimodal reward model that integrates visual and semantic understanding to provide reliable, human-aligned evaluations of generated images. Additionally, based on StyleReward-Dataset, we develop a two-stage training framework for style transfer, consisting of supervised fine-tuning on StyleReward-Dataset followed by reinforcement learning with Group Relative Preference Optimization that converts relative preferences into stable policy updates. Extensive experiments on both public benchmark and proposed benchmark show that our proposed method achieves substantial improvements in style fidelity and content preservation over strong baselines, with human studies further validating its superior visual realism and alignment with human preference.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 64498550-3f1e-4641-bae9-9f9bd5a74948

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines