Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
Bowei Chen, Sai Bi, Hao Tan, He Zhang, Tianyuan Zhang, Zhengqi Li, Yuanjun Xiong, Jianming Zhang, Kai Zhang
摘要
In this work, we propose aligning pretrained visual encoders to serve as tokenizers for latent diffusion models in image generation. Unlike training a variational autoencoder (VAE) from scratch, which primarily emphasizes low-level details, our approach leverages the rich semantic structure of foundation encoders. We introduce a three-stage alignment strategy called AlignTok: (1) freeze the encoder and train an adapter and a decoder to establish a semantic latent space; (2) jointly optimize all components with an additional semantic preservation loss, enabling the encoder to capture perceptual details while retaining high-level semantics; and (3) refine the decoder for improved reconstruction quality. This alignment yields semantically rich image tokenizers that benefit diffusion models. On ImageNet 256256, our tokenizer accelerates the convergence of diffusion models, reaching a gFID of 1.90 within just 64 epochs, and improves generation both with and without classifier-free guidance. Scaling to LAION, text-to-image models trained with our tokenizer consistently outperforms FLUX VAE and VA-VAE under the same training steps. Overall, our method is simple, scalable, and establishes a semantically grounded paradigm for continuous tokenizer design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and EditingShilong Zhang, He Zhang, Zhifei Zhang, Chongjian GE 等ICML 2026 · 被引用 19 次
- Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video GeneratorHyojun Go, Dominik Narnhofer, Goutam Bhat, Prune Truong 等ICLR 2026 · 被引用 9 次
- RecTok: Reconstruction Distillation along Rectified FlowQingyu Shi, Size Wu, Jinbin Bai, Kaidong Yu 等CVPR 2026 · 被引用 5 次
- DA-VAE: Plug-in Latent Compression for Diffusion via Detail AlignmentXin Cai, Zhiyuan You, Zhoutong Zhang, Tianfan XueCVPR 2026 · 被引用 3 次
- End-to-End Autoregressive Image Generation with 1D Semantic TokenizerWenda Chu, Bingliang Zhang, Jiaqi Han, Yizhuo Li 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion ModelsTianci Bi, Xiaoyi Zhang, Yan Lu, Nanning ZhengCVPR 2026 · 被引用 17 次
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsJingfeng Yao, Bin Yang, Xinggang WangCVPR 2025
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive ModelPingyu Wu, Kai Zhu, Yu Liu, Longxiang Tang 等ICLR 2026 · 被引用 16 次
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion TransformersXingjian Leng, Jaskirat Singh, Yunzhong Hou, Zhenchang Xing 等ICCV 2025 · 被引用 15 次
- Vision Foundation Models as Effective Visual Tokenizers for Autoregressive GenerationAnlin Zheng, Xin Wen, Xuanyang Zhang, Chuofan Ma 等NeurIPS 2025 · 被引用 21 次
