Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
Yi Wu, Shengju Qian, Lingting Zhu, Lei Liu, Wandi Qiao, Ziqiang Li, Lequan Yu, Bin Li
Abstract
Multimodal autoregressive (AR) models, based on nexttoken prediction and transformer architecture, have demonstrated remarkable capabilities in various multimodal tasks including text-to-image (T2I) generation. Despite their strong performance in general T2I tasks, our research reveals that these models initially struggle with subject-driven image generation compared to dominant diffusion models. To address this limitation, we introduce Proxy-Tuning, leveraging diffusion models to enhance AR models' capabilities in subject-specific image generation. Our method reveals a striking weak-to-strong phenomenon: fine-tuned AR models consistently outperform their diffusion model supervisors in both subject fidelity and prompt adherence. We analyze this performance shift and identify scenarios where AR models excel, particularly in multi-subject compositions and contextual understanding. This work not only demonstrates impressive results in subject-driven AR image generation, but also unveils the potential of weak-to-strong generalization in the image generation domain, contributing to a deeper understanding of different architectures' strengths and limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad7317a2-c442-421f-94b8-70e489d8985cCited by top-tier papers3
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationAbdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter WonkaNeurIPS 2025 · 6 citations
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image GenerationQian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei et al.AAAI 2026 · 6 citations
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image GenerationFangtai Wu, Mushui Liu, Weijie He, Zhao Wang et al.CVPR 2026 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- Fine-Tuning Visual Autoregressive Models for Subject-Driven GenerationJiwoo Chung, Sangeek Hyun, Hyunjun Kim, Eunseo Koh et al.ICCV 2025 · 1 citation
- UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image GenerationLunhao Duan, Shanshan Zhao, Wenjun Yan, Yinglun Li et al.CVPR 2025
- Hierarchical Masked Autoregressive Models with Low-Resolution Token PivotsGuangting Zheng, Yehao Li, Yingwei Pan, Jiajun Deng et al.ICML 2025
- From Prediction to Perfection: Introducing Refinement to Autoregressive Image GenerationCheng Cheng, Lin Song, Di An, Yicheng Xiao et al.ICLR 2026 · 3 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
