Dr2Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
Chen Zhao, Shuming Liu, Karttikeya Mangalam, Guocheng Qian, Fatimah Zohra, Abdulmohsen Alghannam, Jitendra Malik, Bernard Ghanem
摘要
Large pretrained models are increasingly crucial in modern computer vision tasks. These models are typically used in downstream tasks by end-to-end finetuning, which is highly memory-intensive for tasks with high-resolution data, e.g., video understanding, small object detection, and point cloud analysis. In this paper, we propose Dynamic Reversible Dual-Residual Networks, or Dr 2 Net, a novel family of network architectures that acts as a surrogate network to finetune a pretrained model with substantially reduced memory consumption. Dr 2 Net contains two types of residual connections, one maintaining the residual structure in the pretrained models, and the other making the network reversible. Due to its reversibility, intermediate activations, which can be reconstructed from output, are cleared from memory during training. We use two coefficients on either type of residual connections respectively, and introduce a dynamic training strategy that seamlessly transitions the pretrained model to a reversible network with much higher numerical precision. We evaluate Dr 2 Net on various pretrained models and various tasks, and show that it can reach comparable performance to conventional finetuning but with significantly less memory usage. Code will be available at https://github.com/ coolbay/Dr2Net.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- End-to-End Temporal Action Detection with 1B Parameters Across 1000 FramesShuming Liu, Chen-Lin Zhang, Chen Zhao, Bernard GhanemCVPR 2024 · 被引用 35 次
- TokenSeek: Memory Efficient Fine Tuning via Instance-Aware Token DitchingRunjia Zeng, Qifan Wang, Qiang Guan, Ruixiang Tang 等ICLR 2026 · 被引用 1 次
- NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationQi Bi, Jingjun Yi, Huimin Huang, Hao Zheng 等CVPR 2025
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- Efficient and Information-Preserving Future Frame Prediction and BeyondWei Yu, Yichao Lu, Steve Easterbrook, Sanja FidlerICLR 2020 · 被引用 127 次
- Re2TAL: Rewiring Pretrained Video Backbones for Reversible Temporal Action LocalizationChen Zhao, Shuming Liu, Karttikeya Mangalam, Bernard GhanemCVPR 2023
- Reversible Vision TransformersKarttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu 等CVPR 2022 · 被引用 52 次
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang 等NeurIPS 2021 · 被引用 71 次
- Deep Reorganization: Retaining Residuals in TinyMLHashan Roshantha Mendis, Chih-Kai Kang, Chun-Han Lin, Ming-Syan Chen 等DAC 2024 · 被引用 1 次
