Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-Experts
Tao Zhong, Zhixiang Chi, Li Gu, Yang Wang, Yuanhao Yu, Jin Tang
摘要
In this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not adapted. Furthermore, expecting single-model training to learn extensive knowledge from multiple source domains is counterintuitive. The model is more biased toward learning only domain-invariant features and may result in negative knowledge transfer. In this work, we propose a novel framework for unsupervised test-time adaptation, which is formulated as a knowledge distillation process to address domain shift. Specifically, we incorporate Mixture-of-Experts (MoE) as teachers, where each expert is separately trained on different source domains to maximize their specialty. Given a test-time target domain, a small set of unlabeled data is sampled to query the knowledge from MoE. As the source domains are correlated to the target domains, a transformer-based aggregator then combines the domain knowledge by examining the interconnection among them. The output is treated as a supervision signal to adapt a student prediction network toward the target domain. We further employ meta-learning to enforce the aggregator to distill positive knowledge and the student network to achieve fast adaptation. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art and validates the effectiveness of each proposed component. Our code is available at https://github.com/n3il666/Meta-DMoE . * equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- MetaGCD: Learning to Continually Learn in Generalized Category DiscoveryYanan Wu, Zhixiang Chi, Yang Wang, Songhe FengICCV 2023 · 被引用 50 次
- Test-Time Domain Adaptation by Learning Domain-Aware Batch NormalizationYanan Wu, Zhixiang Chi, Yang Wang, Konstantinos N. Plataniotis 等AAAI 2024 · 被引用 41 次
- BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time AdaptationDaeun Lee, Jaehong Yoon, Sung Ju HwangICML 2024 · 被引用 27 次
- Adapting to Distribution Shift by Visual Domain Prompt GenerationZhixiang Chi, Li Gu, Tao Zhong, Huan Liu 等ICLR 2024 · 被引用 23 次
- AdapTraj: A Multi-Source Domain Generalization Framework for Multi-Agent Trajectory PredictionTangwen Qian, Yile Chen, Gao Cong, Yongjun Xu 等ICDE 2024 · 被引用 15 次
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
相关 Paper
- Mixture of Prototypes for Test-time Adaptive SegmentationGuangrui Li, Zhengyu Zhu, Yongxin GeCVPR 2026 · 被引用 1 次
- Test-Time Adaptation of Medical Vision-Language Models with Mixture of Modality ExpertsHancong Wang, Yue Yu, Hairong Zheng, Tong ZhangACM MM 2025
- Transformer Based Multi-Source Domain AdaptationDustin Wright, Isabelle AugensteinEMNLP 2020 · 被引用 3 次
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 被引用 5 次
- MoETTA: Test-Time Adaptation Under Mixed Distribution Shifts with MoE-LayerNormXiao Fan, Jingyan Jiang, Zhaoru Chen, Fanding Huang 等AAAI 2026
