RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
Zhengjia Zhong, Shuyan Ke, Zaizhou Lin, Jiaqi Song, Hongyi Lan, Hui Li
Abstract
Vector quantization is a fundamental tool for compressing high-dimensional embeddings, yet existing multi-codebook methods rely on static codebooks that limit expressiveness under heterogeneous data geometry. While recent dynamic quantizers like QINCo adapt codebooks to individual inputs and improve expressiveness, their strict sequential dependencies create decoding bottlenecks. We propose Residual Quantization via Mixture of Experts (RQ-MoE), a framework combining a two-level MoE with dual-stream quantization to enable input-dependent codebook adaptation for efficient vector quantization. RQ-MoE enables dynamic codebook construction and decouples instruction from quantization, facilitating parallel decoding. Theoretically, we show that standard Residual Quantization and QINCo can be recovered as constrained special cases of RQ-MoE, and derive a guideline for setting expert dimensionality in RQ-MoE. Extensive experiments show that RQ-MoE achieves state-ofthe-art or on-par performance in reconstruction and retrieval, while it can provide 6×-14× faster decoding than prior vector quantization methods. The implementation is available at https: //github.com/KDEGroup/RQ-MoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho et al.CVPR 2022 · 184 citations
- Unsupervised Neural Quantization for Compressed-Domain Similarity SearchStanislav Morozov, Artem BabenkoICCV 2019 · 31 citations
- Residual Quantization with Implicit Neural CodebooksIris A. M. Huijben, Matthijs Douze, Matthew J. Muckley, Ruud van Sloun et al.ICML 2024 · 23 citations
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao et al.VLDB 2024 · 20 citations
Related papers
- Qinco2: Vector Compression and Search with Improved Implicit Neural CodebooksThéophane Vallaeys, Matthew J. Muckley, Jakob Verbeek, Matthijs DouzeICLR 2025
- KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language ModelsZukang Xu, Zhixiong Zhao, Xing Hu, Zhixuan Chen et al.ICLR 2026 · 7 citations
- Delta Decompression for MoE-based LLMs CompressionHao Gu, Wei Li, Lujun Li, Qiyuan Zhu et al.ICML 2025
- D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingHaodong Wang, Qihua Zhou, Zicong Hong, Song GuoMobiCom 2025 · 8 citations
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 67 citations
