In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data Shuffle
Lijie Xu, Shuang Qiu, Binhang Yuan, Jiawei Jiang, Cédric Renggli, Shaoduo Gan, Kaan Kara, Guoliang Li, Ji Liu, Wentao Wu, Jieping Ye, Ce Zhang
Abstract
Stochastic gradient descent (SGD) is the cornerstone of modern ML systems. Despite its computational efficiency, SGD requires random data access that is inherently inefficient when implemented in systems that rely on block-addressable secondary storage such as HDD and SSD, e.g., in-DB ML systems and TensorFlow/PyTorch over large files. To address this impedance mismatch, various data shuffling strategies have been proposed to balance the convergence rate of SGD (which favors randomness) and its I/O performance (which favors sequential access).
In this paper, we first conduct a systematic empirical study on existing data shuffling strategies, which reveals that all existing strategies have room for improvement-they suffer in terms of I/O performance or convergence rate. With this in mind, we propose a simple but novel hierarchical data shuffling strategy, CorgiPile. Compared with existing strategies, CorgiPile avoids a full data shuffle while maintaining comparable convergence rate of SGD as if a full shuffle were performed. We provide a non-trivial theoretical analysis of CorgiPile on its convergence behavior. We further integrate CorgiPile into PostgreSQL by introducing three new physical operators with optimizations. Our experimental results show that CorgiPile can achieve comparable convergence rate to the full shuffle based SGD, and 1.6×-12.8× faster than two state-ofthe-art in-DB ML systems, Apache MADlib and Bismarck, on both HDD and SSD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local UpdateFangcheng Fu, Xupeng Miao, Jiawei Jiang, Huanran Xue et al.VLDB 2022 · 31 citations
- Database Native Model Selection: Harnessing Deep Neural Networks in Database SystemsNaili Xing, Shaofeng Cai, Gang Chen, Zhaojing Luo et al.VLDB 2024 · 14 citations
- Auto-Differentiation of Relational Computations for Very Large Scale Machine LearningYuxin Tang, Zhimin Ding, Dimitrije Jankov, Binhang Yuan et al.ICML 2023 · 7 citations
- Powering In-Database Dynamic Model Slicing for Structured Data AnalyticsLingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen et al.VLDB 2024 · 7 citations
- GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System ResearchMeng Wang, Gus Waldspurger, Naufal Ananda, Yuyang Huang et al.VLDB 2025 · 1 citation
Builds on9
- Random Reshuffling: Simple Analysis with Vast ImprovementsKonstantin Mishchenko, Ahmed Khaled, Peter RichtárikNeurIPS 2020 · 172 citations
- Closing the convergence gap of SGD without replacementShashank Rajput, Anant Gupta, Dimitris S. PapailiopoulosICML 2020 · 73 citations
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- Tensor Relational Algebra for Distributed Machine Learning System DesignBinhang Yuan, Dimitrije Jankov, Jia Zou, Yuxin Tang et al.VLDB 2021 · 33 citations
Related papers
- The benefits of full data shuffle, now with optimal I/O cost: -wise independence and matrix transposition to the rescuePeyman Afshani, Rezaul Chowdhury, Mayank Goswami, Jens Kristian R Schou et al.ICML 2026
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- C olumnSGD: A Column-oriented Framework for Distributed Stochastic Gradient DescentZhipeng Zhang, Wentao Wu, Jiawei Jiang, Lele Yu et al.ICDE 2020 · 6 citations
- Tighter Convergence Bounds for Shuffled SGD via Primal-Dual PerspectiveXufeng Cai, Cheuk Yin Lin, Jelena DiakonikolasNeurIPS 2024 · 9 citations
- GraB: Finding Provably Better Data Permutations than Random ReshufflingYucheng Lu, Wentao Guo, Christopher De SaNeurIPS 2022 · 23 citations
