Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
Fred Hohman, Chaoqun Wang, Jinmook Lee, Jochen Görtler, Dominik Moritz, Jeffrey P. Bigham, Zhile Ren, Cecile Foret, Qi Shan, Xiaoyi Zhang
摘要
On-device machine learning (ML) moves computation from the cloud to personal devices, protecting user privacy and enabling intelligent user experiences. However, fitting models on devices with limited resources presents a major technical challenge: practitioners need to optimize models and balance hardware metrics such as model size, latency, and power. To help practitioners create efficient ML models, we designed and developed Talaria : a model visualization and optimization system. Talaria enables practitioners to compile models to hardware, interactively visualize model statistics, and simulate optimizations to test the impact on inference metrics. Since its internal deployment two years ago, we have evaluated Talaria using three methodologies: (1) a log analysis highlighting its growth of 800+ practitioners submitting 3,600+ models; (2) a usability survey with 26 users assessing the utility of 20 Talaria features; and (3) a qualitative interview with the 7 most active users about their experience using Talaria.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression ExperimentsAngie W. Boggust, Venkatesh Sivaraman, Yannick Assogba, Donghao Ren 等IEEE VIS 2024 · 被引用 12 次
- Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual RelationshipsAngie W. Boggust, Hyemin Bang, Hendrik Strobelt, Arvind SatyanarayanCHI 2025 · 被引用 4 次
- Qua2SeDiMo: Quantifiable Quantization Sensitivity of Diffusion ModelsKeith G. Mills, Mohammad Salameh, Ruichen Chen, Negar Hassanpour 等AAAI 2025
它引用的顶会 Paper13
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel 等ICCV 2023 · 被引用 341 次
- How do Data Science Workers Collaborate? Roles, Workflows, and ToolsAmy X. Zhang, Michael J. Muller, Dakuo WangCSCW 2020 · 被引用 260 次
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover 等IEEE VIS 2022 · 被引用 191 次
- VATLD: A Visual Analytics System to Assess, Understand and Improve Traffic Light DetectionLiang Gou, Lincan Zou, Nanxiang Li, Michael Hofmann 等IEEE VIS 2020 · 被引用 72 次
- Dynamic Resolution NetworkMingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang 等NeurIPS 2021 · 被引用 71 次
相关 Paper
- Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning ExperiencesFred Hohman, Mary Beth Kery, Donghao Ren, Dominik MoritzCHI 2024 · 被引用 27 次
- MyML: User-Driven Machine LearningVidushi Goyal, Valeria Bertacco, Reetuparna DasDAC 2021 · 被引用 3 次
- A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareYanchi Dong, Tianyu Jia, Kaixuan Du, Yiqi Jing 等DAC 2023 · 被引用 10 次
- Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language ModelsChung-ju Huang, Huiqiang Zhao, Yuanpeng He, Lijian Li 等ACL 2026
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 等USENIX ATC 2025 · 被引用 41 次
