Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
Fred Hohman, Chaoqun Wang, Jinmook Lee, Jochen Görtler, Dominik Moritz, Jeffrey P. Bigham, Zhile Ren, Cecile Foret, Qi Shan, Xiaoyi Zhang
Abstract
On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria : a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression ExperimentsAngie W. Boggust, Venkatesh Sivaraman, Yannick Assogba, Donghao Ren et al.IEEE VIS 2024 · 12 citations
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 4 citations
- Qua2SeDiMo: Quantifiable Quantization Sensitivity of Diffusion ModelsKeith G. Mills, Mohammad Salameh, Ruichen Chen, Negar Hassanpour et al.AAAI 2025
Builds on13
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel et al.ICCV 2023 · 341 citations
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 260 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
- VATLD: A Visual Analytics System to Assess, Understand and Improve Traffic Light DetectionLiang Gou, Lincan Zou, Nanxiang Li, Michael Hofmann et al.IEEE VIS 2020 · 72 citations
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang et al.NeurIPS 2021 · 71 citations
Related papers
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning ExperiencesFred Hohman, Mary Beth Kery, Donghao Ren, Dominik MoritzCHI 2024 · 27 citations
- MyML: User-Driven Machine LearningVidushi Goyal, Valeria Bertacco, Reetuparna DasDAC 2021 · 3 citations
- A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareYanchi Dong, Tianyu Jia, Kaixuan Du, Yiqi Jing et al.DAC 2023 · 10 citations
- Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language ModelsChung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li et al.ACL 2026
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li et al.USENIX ATC 2025 · 41 citations
