On Continual Model Refinement in Out-of-Distribution Data Streams
Bill Yuchen Lin, Sida Wang, Xi Victoria Lin, Robin Jia, Lin Xiao, Xiang Ren, Scott Yih
Abstract
Real-world natural language processing (NLP) models need to be continually updated to fix the prediction errors in out-of-distribution (OOD) data streams while overcoming catastrophic forgetting. However, existing continual learning (CL) problem setups cannot cover such a realistic and complex scenario. In response to this, we propose a new CL problem formulation dubbed continual model refinement (CMR). Compared to prior CL settings, CMR is more practical and introduces unique challenges (boundary-agnostic and non-stationary distribution shift, diverse mixtures of multiple OOD data clusters, errorcentric streams, etc.). We extend several existing CL approaches to the CMR setting and evaluate them extensively. For benchmarking and analysis, we propose a general sampling algorithm to obtain dynamic OOD data streams with controllable non-stationarity, as well as a suite of metrics measuring various aspects of online performance. Our experiments and detailed analysis reveal the promise and challenges of the CMR problem, supporting that studying CMR in dynamic OOD streams can benefit the longevity of deployed NLP models in production. 1 * The work was done when Bill was an intern at FAIR. 1 Our code and data are available at the project website - https://cmr-nlp.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Aging with GRACE: Lifelong Model Editing with Discrete Key-Value AdaptorsTom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim et al.NeurIPS 2023 · 349 citations
- WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language ModelsPeng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu et al.NeurIPS 2024 · 125 citations
- MELO: Enhancing Model Editing with Neuron-Indexed Dynamic LoRALang Yu, Qin Chen, Jie Zhou, Liang HeAAAI 2024 · 96 citations
- Lifelong Language Pretraining with Distribution-Specialized ExpertsWuyang Chen, Yanqi Zhou, Nan Du, Yanping Huang et al.ICML 2023 · 85 citations
- Fine-tuned Language Models are Continual LearnersThomas Scialom, Tuhin Chakrabarty, Smaranda MuresanEMNLP 2022 · 46 citations
Builds on8
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn et al.ICLR 2022 · 527 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
- Towards Continual Knowledge Learning of Language ModelsJoel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin et al.ICLR 2022 · 204 citations
Related papers
- Online Fast Adaptation and Knowledge Accumulation (OSAKA): a New Approach to Continual LearningMassimo Caccia, Pau Rodríguez, Oleksiy Ostapenko, Fabrice Normandin et al.NeurIPS 2020 · 83 citations
- Online Continual Learning on Class Incremental Blurry Task Configuration with Anytime InferenceHyunseo Koh, Dahyun Kim, Jung-Woo Ha, Jonghyun ChoiICLR 2022 · 84 citations
- Computationally Budgeted Continual Learning: What Does Matter?Ameya Prabhu, Hasan Abed Al Kader Hammoud, Puneet K. Dokania, Philip H. S. Torr et al.CVPR 2023
- C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal DataHuaiwen Zhang, Yang Yang, Fan Qi, Shengsheng Qian et al.ACM MM 2023 · 8 citations
- Real-Time Evaluation in Online Continual Learning: A New HopeYasir Ghunaim, Adel Bibi, Kumail Alhamoud, Motasem Alfarra et al.CVPR 2023
