Quantized Model Soup Shake-Up: Weight Perturbation for Enhanced Ensemble Diversity
Jinwoo Chung, Sungyeop Jung, Weronika Czorapinska, Jangho Kim
Abstract
Model soup, averaging the weights of multiple fine-tuned models, delivers ensemble-level accuracy at single-model inference cost, but its success requires both linear mode connectivity (LMC) and sufficient diversity among candidates. We study these two requirements under quantization-aware training (QAT). First, we show analytically and empirically that fine-tuning from a sufficiently converged QAT checkpoint, which we call the QAT anchor state, preserves linear mode connectivity by keeping models within the same loss basin. Second, we identify Ensemble Degeneracy, where the many-to-one mapping of the quantizer collapses independently fine-tuned models into nearly identical quantized representations, eliminating diversity despite intact connectivity. To resolve this, we propose Quantized Model Soup Shake-Up (QMSS), which selectively perturbs low-magnitude weights near quantization bin boundaries to flip their integer assignments, then briefly re-trains each variant via QAT. Because only a small fraction of inherently unstable weights are modified, QMSS induces substantial quantized-domain diversity while preserving LMC. Experiments on CIFAR-100, Tiny-ImageNet, ImageNet, and FSD-Kaggle2018 with both EWGS and LSQ quantizers on CNNs and Vision Transformers show consistent improvements over standard model soup, with gains up to +1.54% Top-1 on CIFAR-100 and +2.09% on CIFAR-100-C, alongside over 6.37× inference speedup on a Jetson Orin Nano at 8/8-bit.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Allowing Oscillation Quantization: Overcoming Solution Space Limitation in Low Bit-Width QuantizationWeiying Xie, Zihan Meng, Jitao Ma, Wenjin Guo et al.ICCV 2025 · 1 citation
- MODEL SOUPS NEED ONLY ONE INGREDIENTAlireza Abdollahpourrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal FrossardICML 2026
- Double Rounding: Nearly Lossless Adaptive Bit Switching for QATHaiduo Huang, Zhenhua Liu, Tian Xia, Pengju RenAAAI 2026
- Quantized Feature Distillation for Network QuantizationKe Zhu, Yin-Yin He, Jianxin WuAAAI 2023 · 21 citations
- PARQ: Piecewise-Affine Regularized QuantizationLisa Jin, Jianhao Ma, Zechun Liu, Andrey Gromov et al.ICML 2025
