Serving Deep Learning Models with Deduplication from Relational Databases
Lixi Zhou, Jiaqing Chen, Amitabh Das, Hong Min, Lei Yu, Ming Zhao, Jia Zou
Abstract
Serving deep learning models from relational databases brings significant benefits. First, features extracted from databases do not need to be transferred to any decoupled deep learning systems for inferences, and thus the system management overhead can be significantly reduced. Second, in a relational database, data management along the storage hierarchy is fully integrated with query processing, and thus it can continue model serving even if the working set size exceeds the available memory. Applying model deduplication can greatly reduce the storage space, memory footprint, cache misses, and inference latency. However, existing data deduplication techniques are not applicable to the deep learning model serving applications in relational databases. They do not consider the impacts on model inference accuracy as well as the inconsistency between tensor blocks and database pages. This work proposed synergistic storage optimization techniques for duplication detection, page packing, and caching, to enhance database systems for model serving. Evaluation results show that our proposed techniques significantly improved the storage efficiency and the model inference latency, and outperformed existing deep learning frameworks in targeting scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a136781-e205-4ffd-82a1-59e3cf0ef111Cited by top-tier papers8
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li et al.SIGMOD 2026 · 24 citations
- Experimental Analysis of Large-scale Learnable Vector Storage CompressionHailin Zhang, Penghao Zhao, Xupeng Miao, Yingxia Shao et al.VLDB 2024 · 20 citations
- Auto-Differentiation of Relational Computations for Very Large Scale Machine LearningYuxin Tang, Zhimin Ding, Dimitrije Jankov, Binhang Yuan et al.ICML 2023 · 7 citations
- Naive Bayes Classifiers over Missing Data: Decision and PoisoningSong Bian, Xiating Ouyang, Zhiwei Fan, Paraschos KoutrisICML 2024 · 5 citations
- Privacy and Accuracy-Aware AI/ML Model DeduplicationHong Guan, Lei Yu, Lixi Zhou, Li Xiong et al.SIGMOD 2025 · 3 citations
Builds on6
- A Tensor Compiler for Unified Machine Learning Prediction ServingSupun Nakandala, Karla Saur, Gyeong-In Yu, Konstantinos Karanasos et al.OSDI 2020 · 60 citations
- Tensors: An abstraction for general data processingDimitrios Koutsoukos, Supun Nakandala, Konstantinos Karanasos, Karla Saur et al.VLDB 2021 · 38 citations
- Austere Flash Caching with Deduplication and CompressionQiuping Wang, Jinhong Li, Wen Xia, Erik Kruus et al.USENIX ATC 2020 · 26 citations
- A Relational Matrix Algebra and its Implementation in a Column StoreOksana Dolmatova, Nikolaus Augsten, Michael H. BöhlenSIGMOD 2020 · 11 citations
- Lachesis: Automated Partitioning for UDF-Centric AnalyticsJia Zou, Amitabh Das, Pratik Barhate, Arun Iyengar et al.VLDB 2021 · 1 citation
Related papers
- Tetris: Memory-efficient Serverless Inference through Tensor SharingJie Li, Laiping Zhao, Yanan Yang, Kunlin Zhan et al.USENIX ATC 2022
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai et al.VLDB 2024 · 17 citations
- Beyond Inference: Performance Analysis of DNN Server Overheads for Computer VisionAhmed F. AbouElhamayed, Susanne Balle, Deshanand P. Singh, Mohamed S. AbdelfattahDAC 2024 · 3 citations
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang et al.INFOCOM 2025 · 2 citations
