ChameleonAPI: Automatic and Efficient Customization of Neural Networks for ML Applications
Yuhan Liu, Chengcheng Wan, Kuntai Du, Henry Hoffmann, Junchen Jiang, Shan Lu, Michael Maire
Abstract
ML APIs have greatly relieved application developers of the burden to design and train their own neural network models-classifying objects in an image can now be as simple as one line of Python code to call an API. However, these APIs offer the same pre-trained models regardless of how their output is used by different applications. This can be suboptimal as not all ML inference errors can cause application failures, and the distinction between inference errors that can or cannot cause failures varies greatly across applications.
To tackle this problem, we first study 77 real-world applications, which collectively use six ML APIs from two providers, to reveal common patterns of how ML API output affects applications' decision processes. Inspired by the findings, we propose ChameleonAPI, an optimization framework for ML APIs, which takes effect without changing the application source code. ChameleonAPI provides application developers with a parser that automatically analyzes the application to produce an abstract of its decision process, which is then used to devise an application-specific loss function that only penalizes API output errors critical to the application. Chameleon-API uses the loss function to efficiently train a neural network model customized for each application and deploys it to serve API invocations from the respective application via existing interface. Compared to a baseline that selects the best-of-all commercial ML API, we show that ChameleonAPI reduces incorrect application decisions by 43%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5a4a43f-92c3-4ef4-be58-01a5824bebb1Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy et al.ICCV 2021 · 778 citations
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau et al.USENIX ATC 2020 · 398 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
Related papers
- Are Machine Learning Cloud APIs Used Correctly?Chengcheng Wan, Shicheng Liu, Henry Hoffmann, Michael Maire et al.ICSE 2021 · 37 citations
- Run-Time Prevention of Software Integration Failures of Machine Learning APIsChengcheng Wan, Yuhan Liu, Kuntai Du, Henry Hoffmann et al.OOPSLA 2023 · 5 citations
- FrugalML: How to use ML Prediction APIs more accurately and cheaplyLingjiao Chen, Matei Zaharia, James Y. ZouNeurIPS 2020 · 57 citations
- How Did the Model Change? Efficiently Assessing Machine Learning API ShiftsLingjiao Chen, Matei Zaharia, James ZouICLR 2022 · 5 citations
- Efficient Online ML API Selection for Multi-Label Classification TasksLingjiao Chen, Matei Zaharia, James ZouICML 2022 · 22 citations
