Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
Yuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu, Yang Shi, Xiaochen Ma, Hao Liang, Yuanxing Zhang, Wentao Zhang
Abstract
Subject-driven image generation has advanced from single-to multi-subject composition, while neglecting distinction, the ability to distinguish and generate the correct subject when inputs contain multiple candidates. This limitation restricts effectiveness in complex, realistic visual settings. We propose Scone, a unified understandinggeneration method that integrates composition and distinction. Scone enables the understanding expert to act as a semantic bridge, conveying semantic information and guiding the generation expert to preserve subject identity while minimizing interference. A two-stage training scheme first learns composition, then enhances distinction through semantic alignment and attention-based masking. We also introduce SconeEval, a benchmark for evaluating both composition and distinction across diverse scenarios. Experiments demonstrate that Scone outperforms existing opensource models in composition and distinction tasks on two benchmarks. Our model, benchmark, and training data are available at: https://github.com/Ryann-Ran/Scone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b24b3f0-cfeb-40b0-9b2e-d58449c2e310Cited by top-tier papers2
- Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative ModelsShiran Ge, Chenyi Huang, Yuang Ai, Qihang Fan et al.CVPR 2026 · 8 citations
- Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity RecognitionXuemei Jia, Jiawei Du, Hui Wei, Jun Chen et al.CVPR 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Show-o2: Improved Native Unified Multimodal ModelsJinheng Xie, Zhenheng Yang, Mike Zheng ShouNeurIPS 2025 · 261 citations
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan et al.CVPR 2026 · 231 citations
- Revealing Single Frame Bias for Video-and-Language LearningJie Lei, Tamara L. Berg, Mohit BansalACL 2023 · 68 citations
Related papers
- Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and GenerationJinyu Liu, Xincheng Shuai, Henghui Ding, Yu-Gang JiangICML 2026 · 2 citations
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang et al.CVPR 2026 · 35 citations
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian et al.ACL 2026 · 19 citations
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu et al.ICLR 2026 · 15 citations
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen et al.NeurIPS 2025 · 61 citations
