ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
Tao Feng, Wei Li, Didi Zhu, Hangjie Yuan, Wendi Zheng, Dan Zhang, Jie Tang
摘要
Backpropagation provides a generalized configuration for overcoming catastrophic forgetting. Optimizers such as SGD and Adam are commonly used for weight updates in continual learning and continual pre-training. However, access to gradient information is not always feasible in practice due to black-box APIs, hardware constraints, or non-differentiable systems, a challenge we refer to as the gradient bans. To bridge this gap, we introduce ZeroFlow, the first benchmark designed to evaluate gradient-free optimization algorithms for overcoming forgetting. ZeroFlow examines a suite of forward pass-based methods across various algorithms, forgetting scenarios, and datasets. Our results show that forward passes alone can be sufficient to mitigate forgetting. We uncover novel optimization principles that highlight the potential of forward pass-based methods in mitigating forgetting, managing task conflicts, and reducing memory demands. Additionally, we propose new enhancements that further improve forgetting resistance using only forward passes. This work provides essential tools and insights to advance the development of forward-pass-based methods for continual learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical ProblemsShuhang Chen, Hangjie Yuan, Yunqiu Xu, Pengwei Liu 等ACL 2026 · 被引用 9 次
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningBorui Kang, Lei Wang, Zhiping Wu, Tao Feng 等ICCV 2025 · 被引用 4 次
- A Faster Path to Continual LearningWei Li, Hangjie Yuan, Zixiang Zhao, Borui Kang 等CVPR 2026 · 被引用 2 次
- Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningXueyi Zhang, Chengwei Zhang, Zheng Li, Xiyu Wang 等AAAI 2026 · 被引用 1 次
- Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language ModelsZiwei Liu, Borui Kang, Wei Li, Hangjie Yuan 等AAAI 2026
它引用的顶会 Paper27
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang 等ICML 2022 · 被引用 343 次
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan 等NeurIPS 2021 · 被引用 229 次
相关 Paper
- Forward-Only Continual LearningJiao Chen, Jiayi He, Fangfang Chen, Zuohong Lv 等ACM MM 2025 · 被引用 1 次
- Make Continual Learning Stronger via C-FlatAng Bian, Wei Li, Hangjie Yuan, Chengrong Yu 等NeurIPS 2024 · 被引用 48 次
- Is Forgetting Less a Good Inductive Bias for Forward Transfer?Jiefeng Chen, Timothy Nguyen, Dilan Görür, Arslan ChaudhryICLR 2023 · 被引用 1 次
- FOZO: Forward-Only Zeroth-Order Prompt Optimization for Test-Time AdaptationXingyu Wang, Tao WangCVPR 2026 · 被引用 2 次
- Continual Learning with Recursive Gradient OptimizationHao Liu, Huaping LiuICLR 2022 · 被引用 52 次
