Lune

AAAI2026Top-tier venue

MoE^2: A Mixture-of-Mixtures of Experts for Ensemble-Free Domain Generalization

Ahmed Radwan, Mahmoud Soliman, Omar Abdelaziz, Ahmad Abdel-Qader, Mohamed S. Shehata

2026Year

Abstract

Domain Generalization (DG) requires models to generalize across unseen data distributions. Kernel-based theory reveals a No-Free-Lunch problem: any model with a fixed representation is fundamentally sub-optimal for all possible shifts. While large ensembles mitigate this, they are computationally expensive and remain static once trained, inheriting the same theoretical limitation. We introduce MoE 2 (Mixture-of-Mixtures of Experts), a framework that uses a single frozen backbone to dynamically synthesize a bespoke adapter for each input, allowing it to continuously adapt its effective kernel. We provide a theoretical grounding for this process, proving our routing mechanism is a principled non-parametric estimator for the optimal Bayes mixture of experts. We derive a generalization bound that cleanly separates the router's estimation error from the reduction in a kernel-mismatch penalty achieved via synthesis. MoE 2 matches or exceeds state-ofthe-art ensemble baselines on major DG benchmarks while using only a single, compact model. MoE 2 thus provides a theoretically-grounded and lightweight alternative to largescale ensembles for robust domain generalization.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines