Task-customized Masked Autoencoder via Mixture of Cluster-conditional Experts
Zhili Liu, Kai Chen, Jianhua Han, Lanqing Hong, Hang Xu, Zhenguo Li, James T. Kwok
Abstract
Masked Autoencoder (MAE) is a prevailing self-supervised learning method that achieves promising results in model pre-training. However, when the various downstream tasks have data distributions different from the pre-training data, the semantically irrelevant pre-training information might result in negative transfer, impeding MAE's scalability. To address this issue, we propose a novel MAEbased pre-training paradigm, Mixture of Cluster-conditional Experts (MoCE), which can be trained once but provides customized pre-training models for diverse downstream tasks. Different from the mixture of experts (MoE), our MoCE trains each expert only with semantically relevant images by using cluster-conditional gates. Thus, each downstream task can be allocated to its customized model pretrained with data most similar to the downstream data. Experiments on a collection of 11 downstream tasks show that MoCE outperforms the vanilla MAE by 2.45% on average. It also obtains new state-of-the-art self-supervised learning results on detection and segmentation. * Equal contribution. 1 We refer to ImageNet-1K as ImageNet if not specified in this paper.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65d50584-0cfd-4eea-8531-e0944bc9ca97Cited by top-tier papers10
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong et al.ICLR 2024 · 248 citations
- Customizing Language Models with Instance-wise LoRA for Sequential RecommendationXiaoyu Kong, Jiancan Wu, An Zhang, Leheng Sheng et al.NeurIPS 2024 · 66 citations
- GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data GenerationKai Chen, Enze Xie, Zhe Chen, Yibo Wang et al.ICLR 2024 · 60 citations
- DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and PerceptionYibo Wang, Ruiyuan Gao, Kai Chen, Kaiqiang Zhou et al.CVPR 2024 · 14 citations
- Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong et al.ACL 2025 · 13 citations
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
Related papers
- Self-Guided Masked AutoencoderJeongwoo Shin, Inseo Lee, Junho Lee, Joonseok LeeNeurIPS 2024 · 18 citations
- CLIP-FMoE: Scalable CLIP via Fused Mixture-of-Experts with Enforced SpecializationLuong Tran, Lan-Cuong Nguyen, Huynh Dang Nguyen, Dat Nguyen-Cong et al.ICLR 2026
- Task-Customized Self-Supervised Pre-training with Scalable Dynamic RoutingZhili Liu, Jianhua Han, Lanqing Hong, Hang Xu et al.AAAI 2022 · 30 citations
- Learning Mask Invariant Mutual Information for Masked Image ModelingTao Huang, Yanxiang Ma, Shan You, Chang XuICLR 2025
- Mixed Autoencoder for Self-Supervised Visual Representation LearningKai Chen, Zhili Liu, Lanqing Hong, Hang Xu et al.CVPR 2023
