Parameter-Efficient Masking Networks
Yue Bai, Huan Wang, Xu Ma, Yitian Zhang, Zhiqiang Tao, Yun Fu
摘要
A deeper network structure generally handles more complicated non-linearity and performs more competitively. Nowadays, advanced network designs often contain a large number of repetitive structures (e.g., Transformer). They empower the network capacity to a new level but also increase the model size inevitably, which is unfriendly to either model restoring or transferring. In this study, we are the first to investigate the representative potential of fixed random weights with limited unique values by learning diverse masks and introduce the Parameter-Efficient Masking Networks (PEMN). It also naturally leads to a new paradigm for model compression to diminish the model size. Concretely, motivated by the repetitive structures in modern neural networks, we utilize one random initialized layer, accompanied with different masks, to convey different feature mappings and represent repetitive network modules. Therefore, the model can be expressed as one-layer with a bunch of masks, which significantly reduce the model storage cost. Furthermore, we enhance our strategy by learning masks for a model filled by padding a given random weights vector. In this way, our method can further lower the space complexity, especially for models without many repetitive architectures. We validate the potential of PEMN learning masks on random weights with limited unique values and test its effectiveness for a new compression paradigm based on different network architectures. Code is available at https://github.com/yueb17/PEMN
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuningHaobo Song, Hao Zhao, Soumajit Majumder, Tao LinICLR 2024 · 被引用 11 次
- DIET: Customized Slimming for Incompatible Networks in Sequential RecommendationKairui Fu, Shengyu Zhang, Zheqi Lv, Jingyuan Chen 等KDD 2024 · 被引用 6 次
- Iterative Soft Shrinkage Learning for Efficient Image Super-ResolutionJiamian Wang, Huan Wang, Yulun Zhang, Yun Fu 等ICCV 2023 · 被引用 5 次
- Distributionally Robust Ensemble of Lottery Tickets Towards Calibrated Sparse Network TrainingHitesh Sapkota, Dingrong Wang, Zhiqiang Tao, Qi YuNeurIPS 2023 · 被引用 5 次
- CHORD: Customizing Hybrid-precision On-device Model for Sequential Recommendation with Device-cloud CollaborationTianqi Liu, Kairui Fu, Shengyu Zhang, Wenyan Fan 等ACM MM 2025
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
相关 Paper
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 被引用 33 次
- DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural NetworksAo Ren, Tao Zhang, Yuhao Wang, Sheng Lin 等AAAI 2020 · 被引用 11 次
- Neural Epitome Search for Architecture-Agnostic Network CompressionDaquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang 等ICLR 2020 · 被引用 13 次
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- Data-Free Network Compression via Parametric Non-uniform Mixed Precision QuantizationVladimir Chikin, Mikhail AntiukhCVPR 2022 · 被引用 17 次
