Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
Aoming Liu, Kevin Miller, Venkatesh Saligrama, Kate Saenko, Boqing Gong, Ser-Nam Lim, Bryan A. Plummer
摘要
Temporal Domain Generalization (TDG) aims to generalize across temporal distribution shifts, e.g., lexical change over time. Prior work often addresses this by predicting future model weights. However, full model prediction is prohibitively expensive for even reasonably sized models. Thus, recent methods only predict the classifier layer, limiting generalization by failing to adjust other model components. To address this, we propose Temporal Experts Averaging (TEA), a novel and scalable TDG framework that updates the entire model using weight averaging to maximize generalization potential while minimizing computational costs. Our theoretical analysis guides us to two steps that enhance generalization to future domains. First, we create expert models with functional diversity yet parameter similarity by fine-tuning a domain-agnostic base model on individual temporal domains while constraining weight changes. Second, we optimize the bias-variance tradeoff through adaptive averaging coefficients derived from modeling temporal weight trajectories in a principal component subspace. Expert's contributions are based on their projected proximity to future domains. Extensive experiments across 7 TDG benchmarks, 5 models, and 2 TDG settings shows TEA outperforms prior TDG methods by up to 69% while being up to 60x more efficient 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 被引用 1,416 次
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho 等NeurIPS 2021 · 被引用 630 次
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu 等ICCV 2019 · 被引用 488 次
相关 Paper
- Temporal Domain Generalization with Drift-Aware Dynamic Neural NetworksGuangji Bai, Chen Ling, Liang ZhaoICLR 2023 · 被引用 6 次
- Weight Diffusion for Future: Learn to Generalize in Non-Stationary EnvironmentsMixue Xie, Shuang Li, Binhui Xie, Chi Harold Liu 等NeurIPS 2024 · 被引用 6 次
- CODA: Temporal Domain Generalization via Concept Drift SimulatorChia-Yuan Chang, Yu-Neng Chuang, Zhimeng Jiang, Kwei-Herng Lai 等KDD 2025
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
- Test-Time Classifier Adjustment Module for Model-Agnostic Domain GeneralizationYusuke Iwasawa, Yutaka MatsuoNeurIPS 2021 · 被引用 456 次
