TiME: Test-Time Mixture-of-Experts Routing via Asymmetric CO-Optimal Transport for Continual Test-Time Adaptation
Tianlun Liu, Zhiliang Tian, Zhen Huang, Tianle Liu, Xingzhi Zhou, Feng Liu, Dongsheng Li
摘要
Large language models usually face continuous domain shifts during testing, which degrade performance on unseen shifting domains. So, researchers propose continual test-time adaptation (CTTA) to adapt to evolving testing domains while preserving knowledge of previous domains, making adaptability-stability (A-S) balance. Existing CTTA methods are constrained by dense base models that encode knowledge from all domains into a global model, hardly achieving the A-S balance. We observe that the model sparsity of mixture-of-experts (MoE) models is better for achieving A–S balance than dense models. In CTTA, however, MoE faces difficulty in (1) correctly routing samples from unseen shifting domains and (2) capturing domain-level shifts. In this paper, we propose test-time mixture-of-experts routing (TiME) via asymmetric co-optimal transport (As-COOT): we model MoE routing in CTTA as a test-time allocation problem via COOT. To ensure reliable routing, we propose a semantic space alignment to align sample-expert distributions via bidirectional contrastive learning. To address COOT’s limitations in CTTA, we propose As-COOT, relaxing sample-side constraints while enforcing expert-side constraints to ensure noise robustness and balance expert load. Experiments show TiME outperforms baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Continual Test-Time Domain AdaptationQin Wang, Olga Fink, Luc Van Gool, Dengxin DaiCVPR 2022 · 被引用 383 次
- Hash Layers For Large Sparse ModelsStephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason WestonNeurIPS 2021 · 被引用 316 次
相关 Paper
- BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time AdaptationDaeun Lee, Jaehong Yoon, Sung Ju HwangICML 2024 · 被引用 27 次
- Mixture of Prototypes for Test-time Adaptive SegmentationGuangrui Li, Zhengyu Zhu, Yongxin GeCVPR 2026 · 被引用 1 次
- Rewiring Experts on the Fly: Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert ModelsGuinan Su, Yanwu Yang, Li Shen, Lu Yin 等ICML 2026 · 被引用 3 次
- Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsTao Zhong, Zhixiang Chi, Li Gu, Yang Wang 等NeurIPS 2022 · 被引用 70 次
- Decomposing the Neurons: Activation Sparsity via Mixture of Experts for Continual Test Time AdaptationRongyu Zhang, Aosong Cheng, Yulin Luo, Gaole Dai 等AAAI 2026
