PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
Zining Wang, Jinyang Guo, Ruihao Gong, Yang Yong, Aishan Liu, Yushi Huang, Jiaheng Liu, Xianglong Liu
Abstract
With the increased attention to model efficiency, post-training sparsity (PTS) has become more and more prevalent because of its effectiveness and efficiency. However, there remain questions on better practice of PTS algorithms and the sparsification ability of models, which hinders the further development of this area.Therefore, a benchmark to comprehensively investigate the issues above is urgently needed. In this paper, we propose the first comprehensive post-training sparsity benchmark called PTSBench towards algorithms and models. We benchmark 10+ PTS general-pluggable fine-grained techniques on 3 typical tasks using over 40 off-the-shelf model architectures. Through extensive experiments and analyses, we obtain valuable conclusions and provide several insights from both algorithms and model aspects. Our PTSBench can provide (1) new observations for a better understanding of the PTS algorithms, (2) in-depth and comprehensive evaluations for the sparsification ability of models, and (3) a well-structured and easy-integrate open-source framework. We hope this work will provide illuminating conclusions and advice for future studies of post-training sparsity methods and sparsification-friendly model design. The code for our PTSBench is released at https://github.com/ModelTC/msbench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4403c0e5-0572-480a-982c-efbe0f8bc035Cited by top-tier papers6
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningLingkun Long, Rubing Yang, Yushi Huang, Desheng Hui et al.AAAI 2026 · 8 citations
- LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitChengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong et al.AAAI 2026 · 3 citations
- HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer AccelerationYushi Huang, Zining Wang, Ruihao Gong, Jing Liu et al.ICML 2025
- CMedBench: A Comprehensive Benchmark for Efficient Medical Large Language ModelsShengbo Gao, Jinyang Guo, Lixian Su, Yifu Ding et al.AAAI 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
Related papers
- UniPTS: A Unified Framework for Proficient Post-Training SparsityJingjing Xie, Yuxin Zhang, Mingbao Lin, Zhihang Lin et al.CVPR 2024
- Fast and Controllable Post-training Sparsity: Learning Optimal Sparsity Allocation with Global Constraint in MinutesRuihao Gong, Yang Yong, Zining Wang, Jinyang Guo et al.AAAI 2024 · 8 citations
- Sparsity May Cry: Let Us Fail (Current) Sparse Neural Networks Together!Shiwei Liu, Tianlong Chen, Zhenyu Zhang, Xuxi Chen et al.ICLR 2023 · 1 citation
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionPeijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li et al.ICML 2025
- EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model CompressionKe Xu, Jiaqi Wan, Wenhao Hu, Han Pu et al.KDD 2026
