MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
Tao Wu, Yibo Jiang, Yehao Lu, Zhizhong Wang, Zeyi Huang, Zequn Qin, Xi Li
摘要
Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Context-Learning based methods are limited by their highly coupled training paradigm. These methods attempt to achieve both high subject fidelity and multi-dimensional human preference alignment within a single training stage, relying on a single, indirect reconstruction loss, which is difficult to simultaneously satisfy both these goals. To address this, we propose MultiCrafter, a framework that decouples this task into two distinct training stages. First, in a pre-training stage, we introduce an explicit positional supervision mechanism that effectively resolves attention bleeding and drastically enhances subject fidelity. Second, in a post-training stage, we propose Identity-Preserving Preference Optimization, a novel online reinforcement learning framework. We feature a scoring mechanism to accurately assess multi-subject fidelity based on the Hungarian matching algorithm, which allows the model to optimize for aesthetics and prompt alignment while ensuring subject fidelity achieved in the first stage. Experiments validate that our decoupling framework significantly improves subject fidelity while aligning with human preferences better.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang 等CVPR 2026 · 被引用 59 次
- Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion modelsRuisi Zhao, Haoren Zheng, Zongxin Yang, Hehe Fan 等ICLR 2026 · 被引用 2 次
- Mixture of Style Experts for Diverse Image StylizationShihao Zhu, Ziheng Ouyang, Yijia Kang, Qilong Wang 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov 等ICLR 2024 · 被引用 816 次
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
相关 Paper
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video GenerationPanwang Pan, Jingjing Zhao, Yuchen Lin, Chenguo Lin 等CVPR 2026 · 被引用 5 次
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D SegmentationPanwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin 等NeurIPS 2025
- MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and DisentanglementDong She, Siming Fu, Mushui Liu, Qiaoqiao Jin 等ICLR 2026 · 被引用 13 次
- FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive FocusQiaoqiao Jin, Siming Fu, Dong She, Weinan Jia 等AAAI 2026 · 被引用 1 次
- PrefPaint: Aligning Image Inpainting Diffusion Model with Human PreferenceKendong Liu, Zhiyu Zhu, Chuanhao Li, Hui Liu 等NeurIPS 2024 · 被引用 26 次
