Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences
Fred Hohman, Mary Beth Kery, Donghao Ren, Dominik Moritz
Abstract
On-device machine learning (ML) promises to improve the privacy, responsiveness, and proliferation of new, intelligent user experiences by moving ML computation onto everyday personal devices. However, today’s large ML models must be drastically compressed to run efficiently on-device, a hurtle that requires deep, yet currently niche expertise. To engage the broader human-centered ML community in on-device ML experiences, we present the results from an interview study with 30 experts at Apple that specialize in producing efficient models. We compile tacit knowledge that experts have developed through practical experience with model compression across different hardware platforms. Our findings offer pragmatic considerations missing from prior work, covering the design process, trade-offs, and technical strategies that go into creating efficient models. Finally, we distill design recommendations for tooling to help ease the difficulty of this work and bring on-device ML into to more widespread practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a60c3d7-1083-45b7-bf4c-4680904095b9Cited by top-tier papers4
- OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question AnsweringJiahao Nick Li, Zhuohao Jerry Zhang, Jiaju MaCHI 2025 · 22 citations
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression ExperimentsAngie W. Boggust, Venkatesh Sivaraman, Yannick Assogba, Donghao Ren et al.IEEE VIS 2024 · 12 citations
- Talaria: Interactively Optimizing Machine Learning Models for Efficient InferenceFred Hohman, Chaoqun Wang, Jinmook Lee, Jochen Görtler et al.CHI 2024 · 8 citations
- On the Interaction of Compressibility and Adversarial RobustnessMelih Barsbey, Antônio H. Ribeiro, Umut Simsekli, Tolga BirdalICLR 2026 · 3 citations
Builds on26
- Questioning the AI: Informing Design Practices for Explainable AI User ExperiencesQ. Vera Liao, Daniel M. Gruen, Sarah MillerCHI 2020 · 758 citations
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Co-Designing Checklists to Understand Organizational Challenges and Opportunities around Fairness in AIMichael A. Madaio, Luke Stark, Jennifer Wortman Vaughan, Hanna M. WallachCHI 2020 · 428 citations
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.ICCV 2023 · 341 citations
- CNN Explainer: Learning Convolutional Neural Networks with Interactive VisualizationZijie J. Wang, Robert Turko, Omar Shaikh, Haekyu Park et al.IEEE VIS 2020 · 341 citations
Related papers
- MyML: User-Driven Machine LearningVidushi Goyal, Valeria Bertacco, Reetuparna DasDAC 2021 · 3 citations
- Mind Your Weight(s): A Large-scale Study on Insufficient Machine Learning Model Protection in Mobile AppsZhichuang Sun, Ruimin Sun, Long Lu, Alan MisloveUSENIX Security 2021 · 101 citations
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural NetworksHakbin Kim, Dong-Wan ChoiSIGMOD 2021 · 1 citation
- Stealing Your Data from Compressed Machine Learning ModelsNuo Xu, Qi Liu, Tao Liu, Zihao Liu et al.DAC 2020 · 3 citations
- AutoMC: Automated Model Compression Based on Domain Knowledge and Progressive SearchChunnan Wang, Hongzhi Wang, Xiangyu ShiICDE 2024 · 2 citations
