A Comparative Study on the Impact of Model Compression Techniques on Fairness in Language Models
Krithika Ramesh, Arnav Chavan, Shrey Pandit, Sunayana Sitaram
Abstract
Compression techniques for deep learning have become increasingly popular, particularly in settings where latency and memory constraints are imposed. Several methods, such as pruning, distillation, and quantization, have been adopted for compressing models, each providing distinct advantages. However, existing literature demonstrates that compressing deep learning models could affect their fairness. Our analysis involves a comprehensive evaluation of pruned, distilled, and quantized language models, which we benchmark across a range of intrinsic and extrinsic metrics for measuring bias in text classification. We also investigate the impact of using multilingual models and evaluation measures. Our findings highlight the significance of considering both the pre-trained model and the chosen compression strategy in developing equitable language technologies. The results also indicate that compression strategies can have an adverse effect on fairness measures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01332c9e-f92d-4c4e-a94b-3a04a9d5d968Cited by top-tier papers6
- The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code GenerationXiaoyu Zhang, Juan Zhai, Shiqing Ma, Qingshuang Bao et al.ACL 2025 · 6 citations
- Surgical Feature-Space Decomposition of LLMs: Why, When and How?Arnav Chavan, Nahush Lele, Deepak K. GuptaACL 2024 · 1 citation
- Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware InitializationJunlin He, Yihong Tang, Tong Nie, Guilong Li et al.ICML 2026
- Bias In, Bias Out? Finding Unbiased Subnetworks in Vanilla ModelsIvan Luiz De Moura Matos, Abdel Djalil Sad Saoud, Ekaterina Lakovleva, Vito Paolo Pastore et al.CVPR 2026
- GuidedQuant: Large Language Model Quantization via Exploiting End Loss GuidanceJinuk Kim, Marwa El Halabi, Wonpyo Park, Clemens J. S. Schaefer et al.ICML 2025
Builds on15
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 160 citations
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian et al.EMNLP 2021 · 113 citations
Related papers
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionJunyuan Hong, Jinhao Duan, Chenhui Zhang, Zhangheng Li et al.ICML 2024 · 54 citations
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language ModelsPhillip Rust, Anders SøgaardICML 2023 · 7 citations
- ZipLM: Inference-Aware Structured Pruning of Language ModelsEldar Kurtic, Elias Frantar, Dan AlistarhNeurIPS 2023 · 69 citations
- Accuracy is Not All You NeedAbhinav Dutta, Sanjeev Krishnan, Nipun Kwatra, Ramachandran RamjeeNeurIPS 2024 · 34 citations
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
