BackSlash: Rate Constrained Optimized Training of Large Language Models
Jun Wu, Jiangtao Wen, Yuxing Han
摘要
The rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on ratedistortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60% -90% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80% pruning rates), and enables network simplification for accelerated inference on edge devices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel 等ICLR 2022 · 被引用 162 次
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationZechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P. Xing 等CVPR 2022 · 被引用 108 次
- Dynamic Structure Pruning for Compressing CNNsJun-Hyung Park, Yeachan Kim, Junho Kim, Joon-Young Choi 等AAAI 2023 · 被引用 24 次
相关 Paper
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 被引用 88 次
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block SkippingYu-Chen Lu, Sheng-Feng Yu, Hui-Hsien Weng, Pei-Shuo Wang 等AAAI 2026 · 被引用 1 次
- SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsRasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker 等ICLR 2025
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu 等DAC 2021 · 被引用 16 次
