STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
Peijie Dong, Lujun Li, Yuedong Zhong, Dayou Du, Ruibo Fan, Yuhan Chen, Zhenheng Tang, Qiang Wang, Wei Xue, Yike Guo, Xiaowen Chu
Abstract
In this paper, we present the first structural binarization method for LLM compression to less than 1-bit precision. Although LLMs have achieved remarkable performance, their memory-bound nature during the inference stage hinders the adoption of resource-constrained devices. Reducing weights to 1-bit precision through binarization substantially enhances computational efficiency. We observe that some weights in binarized LLMs can be randomly flipped without significant performance degradation, suggesting the potential for further compression. To exploit this, our STBLLM employs an N:M sparsity technique to achieve structural binarization of the weights. Specifically, we introduce a novel Standardized Importance (SI) metric, which considers weight magnitude and input feature norm to more accurately assess weight significance. Then, we propose a layer-wise approach, allowing different layers of the LLM to be sparsified with varying N:M ratios, thereby balancing compression and accuracy. Furthermore, we implement a fine-grained grouping strategy for less important weights, applying distinct quantization schemes to sparse, intermediate, and dense regions. Finally, we design a specialized CUDA kernel to support structural binarization. We conduct extensive experiments on LLaMA-1/2/3, OPT family, and Mistral to evaluate the effectiveness of STBLLM. The results demonstrate that our approach performs better than other compressed binarization LLM methods while significantly reducing memory requirements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9baa7c9-b721-4aee-bc84-15b0ee26e80eCited by top-tier papers16
- Quantization Error Propagation: Revisiting Layer-Wise Post-Training QuantizationYamato Arai, Yuma IchikawaNeurIPS 2025 · 46 citations
- QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action ModelsJingxuan Zhang, Yunta Hsieh, Zhongwei Wan, Haokun Lin et al.CVPR 2026 · 24 citations
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language ModelsTianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin et al.ICLR 2026 · 11 citations
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang et al.NSDI 2026 · 8 citations
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- BiLLM: Pushing the Limit of Post-Training Quantization for LLMsWei Huang, Yangdong Liu, Haotong Qin, Ying Li et al.ICML 2024 · 161 citations
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language ModelsWei Huang, Haotong Qin, Yangdong Liu, Yawei Li et al.ICML 2025
- BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary CodebookHao Gu, Lujun Li, Hao Wang, Lei Wang et al.ACL 2026 · 5 citations
- KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV CacheFei Li, Song Liu, Weiguo Wu, Shiqiang Nie et al.AAAI 2026 · 1 citation
- LBLLM: Lightweight Binarization of Large Language Models via Three-Stage DistillationSiqing Song, Chuang Wang, Yong Lang, Yi Yang et al.ACL 2026
