MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis
Tianhong Li, Huiwen Chang, Shlok Kumar Mishra, Han Zhang, Dina Katabi, Dilip Krishnan
摘要
Generative modeling and representation learning are two key tasks in computer vision. However, these models are typically trained independently, which ignores the potential for each task to help the other, and leads to training and model maintenance overheads. In this work, we propose MAsked Generative Encoder (MAGE), the first framework to unify SOTA image generation and self-supervised representation learning. Our key insight is that using variable masking ratios in masked image modeling pre-training can allow generative training (very high masking ratio) and representation learning (lower masking ratio) under the same training framework. Inspired by previous generative models, MAGE uses semantic tokens learned by a vector-quantized GAN at inputs and outputs, combining this with masking. We can further improve the representation by adding a contrastive loss to the encoder output. We extensively evaluate the generation and representation learning capabilities of MAGE. On ImageNet-1K, a single MAGE ViT-L model obtains 9.10 FID in the task of classunconditional image generation and 78.9% top-1 accuracy for linear probing, achieving state-of-the-art performance in both image generation and representation learning. Code is available at https://github.com/LTH14/mage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- 4M: Massively Multimodal Masked ModelingDavid Mizrahi, Roman Bachmann, Oguzhan Fatih Kar, Teresa Yeo 等NeurIPS 2023 · 被引用 154 次
- Machine Unlearning for Image-to-Image Generative ModelsGuihong Li, Hsiang Hsu, Chun-Fu Chen, Radu MarculescuICLR 2024 · 被引用 56 次
- Reinforcement Learning with Maskable Stock Representation for Portfolio Management in Customizable Stock PoolsWentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying 等WWW 2024 · 被引用 18 次
- Image Understanding Makes for A Good Tokenizer for Image GenerationLuting Wang, Yang Zhao, Zijian Zhang, Jiashi Feng 等NeurIPS 2024 · 被引用 16 次
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and QuantizationSiyuan Li, Luyuan Zhang, Zedong Wang, Juanxi Tian 等CVPR 2025
- Learning from Semantic Dictionaries: Discriminative Codebook Contrastive Learning for Unified Visual Representation and GenerationImanol G. Estepa, Jesús M. Rodríguez-de-Vera, Bhalaji Nagarajan, Petia RadevaCVPR 2026
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICCV 2025 · 被引用 3 次
- Masked Auto-Encoders Meet Generative Adversarial Networks and BeyondZhengcong Fei, Mingyuan Fan, Li Zhu, Junshi Huang 等CVPR 2023
