Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
Tianci Bi, Xiaoyi Zhang, Yan Lu, Nanning Zheng
Abstract
The performance of Latent Diffusion Models (LDMs) is critically dependent on the quality of their visual tokenizers. While recent works have explored incorporating Vision Foundation Models (VFMs) into the tokenizers training via distillation, we empirically find this approach inevitably weakens the robustness of learnt representation from original VFM. In this paper, we bypass the distillation by proposing a more direct approach by leveraging the frozen VFM for the LDMs tokenizer, named VFM Variational Autoencoder (VFM-VAE).To fully exploit the potential to leverage frozen VFM for the LDMs tokenizer, we design a new decoder to reconstruct realistic images from the semantic-rich representation of VFM. With the proposed VFM-VAE, we conduct a systematic study on how the representation from different tokenizers impact the representation learning process throughout diffusion training, enabling synergistic benefits of dual-side alignment on both tokenizers and diffusion models. Our effort in tokenizer design and training strategy lead to superior performance and efficiency: our system reaches a gFID (w/o CFG) of 2.22 in merely 80 epochs (a 10 speedup over prior tokenizers). With continued training to 640 epochs, it further attains a gFID (w/o CFG) of 1.62. These results offer solid evidence for the substantial potential of VFMs to serve as visual tokenizers to accelerate the LDM training progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 085b29e0-1f38-4dc8-95b8-3072e73d9e2cCited by top-tier papers5
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency PerspectiveBolin Lai, Xudong Wang, Saketh Rambhatla, James M. Rehg et al.CVPR 2026 · 7 citations
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li et al.ICML 2026 · 2 citations
- RePack then Refine: Efficient Diffusion Transformers with Vision Foundation ModelsGuanfang Dong, Luke Schultz, Negar Hassanpour, Chao GaoICML 2026 · 1 citation
- Flow Matching for Multimodal DistributionsGaoxiang Luo, Frank Cole, Sihang Zhang, Yuxiang Wan et al.CVPR 2026 · 1 citation
- SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion ModelBingliang Zhang, Wenda Chu, Yizhuo Li, Linjie Yang et al.CVPR 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsJingfeng Yao, Bin Yang, Xinggang WangCVPR 2025
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu et al.CVPR 2026 · 5 citations
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive GenerationAnlin Zheng, Xin Wen, Xuanyang Zhang, Chuofan Ma et al.NeurIPS 2025 · 21 citations
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion TransformersXingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing et al.ICCV 2025 · 15 citations
