BackSlash: Rate Constrained Optimized Training of Large Language Models
Jun Wu, Jiangtao Wen, Yuxing Han
Abstract
The rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on ratedistortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60% -90% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80% pruning rates), and enables network simplification for accelerated inference on edge devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel et al.ICLR 2022 · 162 citations
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationZechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P. Xing et al.CVPR 2022 · 108 citations
- Dynamic Structure Pruning for Compressing CNNsJun-Hyung Park, Yeachan Kim, Junho Kim, Joon-Young Choi et al.AAAI 2023 · 24 citations
Related papers
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 88 citations
- SkipCat: Rank-Maximized Low-Rank Compression of Large Language Models via Shared Projection and Block SkippingYu-Chen Lu, Sheng-Feng Yu, Hui-Hsien Weng, Pei-Shuo Wang et al.AAAI 2026 · 1 citation
- SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random GeneratorsRasoul Shafipour, David Harrison, Maxwell Horton, Jeffrey Marker et al.ICLR 2025
- ZeroBN: Learning Compact Neural Networks For Latency-Critical Edge SystemsShuo Huai, Lei Zhang, Di Liu, Weichen Liu et al.DAC 2021 · 16 citations
