Vista: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale
Supun Nakandala, Arun Kumar
Abstract
Scalable systems for machine learning (ML) are largely siloed into dataflow systems for structured data and deep learning systems for unstructured data. This gap has left workloads that jointly analyze both forms of data with poor systems support, leading to both low system efficiency and grunt work for users. We bridge this gap for an important class of such workloads: feature transfer from deep convolutional neural networks (CNNs) for analyzing images along with structured data. Executing feature transfer on scalable dataflow and deep learning systems today faces two key systems issues: inefficiency due to redundant computations and crash-proneness due to mismanaged memory. We present Vista, a new data system that resolves these issues by elevating this workload to a declarative level on top of dataflow and deep learning systems. Vista automatically optimizes the configuration and execution of this workload to reduce both computational redundancy and the potential for workload crashes. Experiments on real datasets show that apart from making feature transfer easier, Vista avoids workload crashes and reduces runtimes by 58% to 92% compared to baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- OmniFair: A Declarative System for Model-Agnostic Group Fairness in Machine LearningHantian Zhang, Xu Chu, Abolfazl Asudeh, Shamkant B. NavatheSIGMOD 2021 · 28 citations
- Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training DatasetsSupun Nakandala, Arun KumarSIGMOD 2022 · 6 citations
Related papers
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun et al.VLDB 2023 · 45 citations
- A Visual Analytics Framework for Explaining and Diagnosing Transfer Learning ProcessesYuxin Ma, Arlen Fan, Jingrui He, Arun Reddy Nelakurthi et al.IEEE VIS 2020 · 36 citations
- Terra: Imperative-Symbolic Co-Execution of Imperative Deep Learning ProgramsTaebum Kim, Eunji Jeong, Geon-Woo Kim, Yunmo Koo et al.NeurIPS 2021 · 7 citations
- HIDA: A Hierarchical Dataflow Compiler for High-Level SynthesisHanchen Ye, Hyegang Jun, Deming ChenASPLOS 2024 · 21 citations
- Beyond Inference: Performance Analysis of DNN Server Overheads for Computer VisionAhmed F. AbouElhamayed, Susanne Balle, Deshanand P. Singh, Mohamed S. AbdelfattahDAC 2024 · 3 citations
