Lune

ICML2026Top-tier venue

NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models

Hyochan Chong, Dongkyu Kim, Changdong Kim, Minseop Choi

2026Year

Abstract

Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binary (1-bit) levels, as they either require large amounts of data and compute or incur additional storage. In this work, we propose NANOQUANT, a post-training quantization (PTQ) method to compress LLMs to both binary and sub-1-bit levels. NANOQUANT formulates quantization as a low-rank binary factorization problem, and compresses full-precision weights to low-rank binary matrices and scales. Specifically, it utilizes an efficient alternating direction method of multipliers (ADMM) solver to precisely initialize latent binary matrices and scales, and then tunes the initialized parameters through a block and model reconstruction process. Consequently, NANOQUANT establishes a new Pareto frontier in low-memory post-training quantization, and enables sub-1-bit compression. NANOQUANT makes large-scale deployment feasible on consumer hardware. For example, it compresses Llama-2-70B by 24× in just 13 hours on a single H100, enabling a 70B model to operate on a consumer 8 GB GPU. Code is available at github.com/SamsungLabs/NanoQuant.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5e6a5b08-fb59-474f-b9f2-4b209ee498c6

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines