Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation
Jiwoo Chung, Sangeek Hyun, Hyunjun Kim, Eunseo Koh, MinKyu Lee, Jae-Pil Heo
Abstract
Recent advances in text-to-image generative models have enabled numerous practical applications, including subject-driven generation, which fine-tunes pretrained models to capture subject semantics from only a few examples. While diffusion-based models produce high-quality images, their extensive denoising steps result in significant computational overhead, limiting real-world applicability. Visual autoregressive (VAR) models, which predict next-scale tokens rather than spatially adjacent ones, offer significantly faster inference suitable for practical deployment. In this paper, we propose the first VAR-based approach for subject-driven generation. However, naive fine-tuning VAR leads to computational overhead, language drift, and reduced diversity. To address these challenges, we introduce selective layer tuning to reduce complexity and prior distillation to mitigate language drift. Additionally, we found that the early stages have a greater influence on the generation of subject than the latter stages, which merely synthesize minor details. Based on this finding, we propose scale-wise weighted tuning, which prioritizes coarser resolutions for promoting the model to focus on the subject-relevant information instead of local details. Extensive experiments validate that our method significantly outperforms diffusion-based baselines across various metrics and demonstrates its practical usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec2707e3-729f-43ae-8acd-44f6cfbf3040Cited by top-tier papers5
- SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion ModelsJiwoo Chung, Sangeek Hyun, MinKyu Lee, Byeongju Han et al.CVPR 2026 · 9 citations
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven GenerationAbdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter WonkaNeurIPS 2025 · 6 citations
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive ModelRuixiao Dong, Zhendong Wang, Keli Liu, Li Li et al.ICLR 2026 · 3 citations
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image GenerationFangtai Wu, Mushui Liu, Weijie He, Zhao Wang et al.CVPR 2026 · 1 citation
- Translation of Text Embedding Via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion ModelsEunseo Koh, Seunghoo Hong, Tae-Young Kim, Simon S. Woo et al.ICCV 2025 · 1 citation
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- RADAR: VQ-VAE Decoder of VAR is a Good Student for Restoring Against Degradation by AccelerationZiyang Wang, Yue Zhang, Mingdao Wang, Yasen Zhang et al.CVPR 2026
- Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image GenerationYi Wu, Shengju Qian, Lingting Zhu, Lei Liu et al.CVPR 2026 · 8 citations
- FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive ModelsSenmao Li, Kai Wang, Salman Khan, Fahad Khan et al.ICML 2026 · 2 citations
- ProtoVAR: Efficient Dataset Distillation via Prototype-Guided Visual Autoregressive ModelingMingyu Wang, Wei JiangICML 2026
- Training-Free Text-Guided Image Editing with Visual Autoregressive ModelYufei Wang, Lanqing Guo, Zhihao Li, Jiaxing Huang et al.ICCV 2025
