MoE^2: A Mixture-of-Mixtures of Experts for Ensemble-Free Domain Generalization
Ahmed Radwan, Mahmoud Soliman, Omar Abdelaziz, Ahmad Abdel-Qader, Mohamed S. Shehata
摘要
Domain Generalization (DG) requires models to generalize across unseen data distributions. Kernel-based theory reveals a No-Free-Lunch problem: any model with a fixed representation is fundamentally sub-optimal for all possible shifts. While large ensembles mitigate this, they are computationally expensive and remain static once trained, inheriting the same theoretical limitation. We introduce MoE 2 (Mixture-of-Mixtures of Experts), a framework that uses a single frozen backbone to dynamically synthesize a bespoke adapter for each input, allowing it to continuously adapt its effective kernel. We provide a theoretical grounding for this process, proving our routing mechanism is a principled non-parametric estimator for the optimal Bayes mixture of experts. We derive a generalization bound that cleanly separates the router's estimation error from the reduction in a kernel-mismatch penalty achieved via synthesis. MoE 2 matches or exceeds state-ofthe-art ensemble baselines on major DG benchmarks while using only a single, compact model. MoE 2 thus provides a theoretically-grounded and lightweight alternative to largescale ensembles for robust domain generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
相关 Paper
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 被引用 5 次
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsQi Wang, Hanyang Peng, Yue YuAAAI 2026 · 被引用 1 次
- Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsTao Zhong, Zhixiang Chi, Li Gu, Yang Wang 等NeurIPS 2022 · 被引用 70 次
- Mixture of Prototypes for Test-time Adaptive SegmentationGuangrui Li, Zhengyu Zhu, Yongxin GeCVPR 2026 · 被引用 1 次
- SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingAthinagoras Skiadopoulos, Mark Zhao, Swapnil Gandhi, Thomas Norrie 等NSDI 2026 · 被引用 6 次
