Learning with Retrospection
Xiang Deng, Zhongfei Zhang
摘要
Deep neural networks have been successfully deployed in various domains of artificial intelligence, including computer vision and natural language processing. We observe that the current standard procedure for training DNNs discards all the learned information in the past epochs except the current learned weights. An interesting question is: is this discarded information indeed useless? We argue that the discarded information can benefit the subsequent training. In this paper, we propose learning with retrospection (LWR) which makes use of the learned information in the past epochs to guide the subsequent training. LWR is a simple yet effective training framework to improve accuracies, calibration, and robustness of DNNs without introducing any additional network parameters or inference cost, but only with a negligible training overhead. Extensive experiments on several benchmark datasets demonstrate the superiority of LWR for training DNNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang 等CVPR 2022 · 被引用 213 次
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 被引用 95 次
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li 等CVPR 2022 · 被引用 88 次
- Comprehensive Knowledge Distillation with Causal InterventionXiang Deng, Zhongfei ZhangNeurIPS 2021 · 被引用 44 次
- A Benchmark Study on CalibrationLinwei Tao, Younan Zhu, Haolan Guo, Minjing Dong 等ICLR 2024 · 被引用 10 次
它引用的顶会 Paper3
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 等CVPR 2020
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
相关 Paper
- Retrospective Loss: Looking Back to Improve Training of Deep Neural NetworksSurgan Jandial, Ayush Chopra, Mausoom Sarkar, Piyush Gupta 等KDD 2020
- Reducing Flipping Errors in Deep Neural NetworksXiang Deng, Yun Xiao, Bo Long, Zhongfei ZhangAAAI 2022 · 被引用 4 次
- Learning to Boost Training by Periodic Nowcasting Near Future WeightsJinhyeok Jang, Woo-han Yun, Won Hwa Kim, Youngwoo Yoon 等ICML 2023 · 被引用 6 次
- Fortuitous Forgetting in Connectionist NetworksHattie Zhou, Ankit Vani, Hugo Larochelle, Aaron C. CourvilleICLR 2022 · 被引用 50 次
- Train-by-Reconnect: Decoupling Locations of Weights from Their ValuesYushi Qiu, Reiji SudaNeurIPS 2020 · 被引用 3 次
