Fast and Scalable Data Transfer Across Data Systems
Haralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker Markl
Abstract
Fast and scalable data transfer is crucial in today's decentralized data ecosystems and data-driven applications. Example use cases include transferring data from operational systems to consolidated data warehouse environments, or from relational database systems to data lakes for exploratory data analysis or ML model training. Traditional data transfer approaches rely on efficient point-to-point connectors or general middleware with generic intermediate data representations. Physical environments (e.g., on-premise, cloud, or consumer nodes) also have become increasingly heterogeneous. Existing work still struggles to achieve both, fast and scalable data transfer as well as generality in terms of heterogeneous systems and environments. Hence, in this paper, we introduce a holistic data transfer framework. Our XDBC framework splits the data transfer pipeline into logical components and provides a wide variety of physical implementations for these components. This design allows a seamless integration of different systems as well as the automatic optimizations of data transfer configurations according to workload and environment characteristics. Our evaluation shows that XDBC outperforms state-of-the-art generic data transfer tools by up to 5x, while being on par with specialized approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0d77269-44dc-4b54-9e94-fea933d7029eCited by top-tier papers1
Ask how each one uses itBuilds on12
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke et al.VLDB 2020 · 109 citations
- Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware OverlaysParas Jain, Sam Kumar, Sarah Wooders, Shishir G. Patil et al.NSDI 2023 · 66 citations
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- The Composable Data Management System ManifestoPedro Pedreira, Orri Erling, Konstantinos Karanasos, Scott Schneider et al.VLDB 2023 · 36 citations
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 32 citations
Related papers
- In-Situ Cross-Database Query ProcessingHaralampos Gavriilidis, Kaustubh Beedkar, Jorge-Arnulfo Quiané-Ruiz, Volker MarklICDE 2023 · 12 citations
- Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science EngineWeizheng Lu, Chao Hui, Yunhai Wang, Feng Zhang et al.VLDB 2025 · 1 citation
- ConnectorX: Accelerating Data Loading From Databases to DataframesXiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen et al.VLDB 2022 · 14 citations
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 7 citations
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 13 citations
