Unified Customized Generation by Disentangled Reward Modeling
Shaojin Wu, Mengqi Huang, Yufeng Cheng, Wenxu Wu, Jiahe Tian, Yiming Luo, Fei Ding, Qian He
Abstract
Existing literature typically treats various customized generation tasks (e.g., subject-customized generation, style-customized generation) as distinct and disjoint problems, with each task focusing solely on customizing a specific aspect of the reference image. However, we argue that the objectives of these different customization tasks are inherently complementary and can be mutually enhanced within a unified framework, as they fundamentally involve the disentanglement of multiple feature aspects from the reference image. To this end, we introduce USO , a U nified S imultaneous O ptimization framework to simultaneously unify different customized tasks (i.e., subject and style). Specifically, USO introduces a cyclical data-model framework that connects these two tasks by a subject-for-style data curation pipeline and a style-for-subject model training pipeline. The subject-for-style data curation pipeline leverages a state-of-the-art subject-customized model to generate high-quality triplet data comprising content images, style images, and their corresponding stylized content images. Building on this foundation, the style-for-subject model training pipeline introduces an auxiliary style reward to simultaneously align style and content features, thereby reinforcing the model’s ability to extract the desired style or content features from the reference image. Extensive experiments demonstrate that USO achieves state-of-the-art performance among open-source models, excelling in both subject consistency and style similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 173fe647-518d-42f0-b92f-59535f3a1653Builds on24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching ParadigmYufeng Cheng, wenxu wu, Shaojin Wu, Mengqi Huang et al.CVPR 2026
- UnZipLoRA: Separating Content and Style from a Single ImageChang Liu, Viraj Shah, Aiyu Cui, Svetlana LazebnikICCV 2025 · 31 citations
- CSGO: Content-Style Composition in Text-to-Image GenerationPeng Xing, Haofan Wang, Yanpeng Sun, Qixun Wang et al.NeurIPS 2025 · 94 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
- UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image GenerationDanning Zhang, Yijing Lin, Shuhan Zhuang, Mengqi Huang et al.ICML 2026
