SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained Environments
Qiuru Lin, Sai Wu, Junbo Zhao, Jian Dai, Meng Shi, Gang Chen, Feifei Li
Abstract
Many IoT applications require the use of multiple deep neural networks (DNNs) to perform various tasks on low-cost edge devices with limited computation resources. However, existing DNN model serving platforms, such as TensorFlow Serving and TorchServe, are resource-intensive and require high-performance GPUs that are often not available on low-cost edge devices. In this paper, we propose SmartLite, a lightweight DBMS that addresses these challenges by storing the parameters and structural information of neural networks as database tables and implementing neural network operators inside the DBMS engine. SmartLite quantizes model parameters as binarized values, applies neural pruning techniques to compress the models, and transforms tensor manipulations into value lookup operations of the DBMS to reduce computation overhead. Experimental results show that SmartLite requires 98% less memory while achieving about a 134% performance speedup compared to Torch-Serve. Our proposed solution addresses the challenges of running multiple DNN models on low-cost edge devices and provides a significant contribution to the field of IoT applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e907635c-77ff-41ca-8dce-1ccbc7f2b459Cited by top-tier papers5
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.SIGMOD 2025 · 6 citations
- Privacy and Accuracy-Aware AI/ML Model DeduplicationHong Guan, Lei Yu, Lixi Zhou, Li Xiong et al.SIGMOD 2025 · 3 citations
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu et al.SIGMOD 2026 · 2 citations
- Alsatian: Optimizing Model Search for Deep Transfer LearningNils Strassenburg, Boris Glavic, Tilmann RablSIGMOD 2025 · 2 citations
- MoDM: Efficient Serving for Image Generation via Mixture-of-Diffusion ModelsYuchen Xia, Divyam Sharma, Yichao Yuan, Souvik Kundu et al.ASPLOS 2026
Builds on12
- MCUNet: Tiny Deep Learning on IoT DevicesJi Lin, Wei-Ming Chen, Yujun Lin, John Cohn et al.NeurIPS 2020 · 827 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Online Knapsack with Frequency PredictionsSungjin Im, Ravi Kumar, Mahshid Montazer Qaem, Manish PurohitNeurIPS 2021 · 70 citations
- Optimizing Machine Learning Inference Queries with Correlative Proxy ModelsZhihui Yang, Zuozhi Wang, Yicong Huang, Yao Lu et al.VLDB 2022 · 33 citations
Related papers
- Memory-Efficient and Secure DNN Inference on TrustZone-enabled Consumer IoT DevicesXueshuo Xie, Haoxu Wang, Zhaolong Jian, Tao Li et al.INFOCOM 2024 · 11 citations
- Serving and Optimizing Machine Learning Workflows on Heterogeneous InfrastructuresYongji Wu, Matthew Lentz, Danyang Zhuo, Yao LuVLDB 2023 · 31 citations
- SmartExchange: Trading Higher-cost Memory Storage/Access for Lower-cost ComputationYang Zhao, Xiaohan Chen, Yue Wang, Chaojian Li et al.ISCA 2020 · 44 citations
- DTMM: Deploying TinyML Models on Extremely Weak IoT Devices with PruningLixiang Han, Zhen Xiao, Zhenjiang LiINFOCOM 2024 · 20 citations
- Mistify: Automating DNN Model Porting for On-Device Inference at the EdgePeizhen Guo, Bo Hu, Wenjun HuNSDI 2021 · 69 citations
