Fast and Scalable Data Transfer Across Data Systems
Haralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker Markl
摘要
Fast and scalable data transfer is crucial in today's decentralized data ecosystems and data-driven applications. Example use cases include transferring data from operational systems to consolidated data warehouse environments, or from relational database systems to data lakes for exploratory data analysis or ML model training. Traditional data transfer approaches rely on efficient point-to-point connectors or general middleware with generic intermediate data representations. Physical environments (e.g., on-premise, cloud, or consumer nodes) also have become increasingly heterogeneous. Existing work still struggles to achieve both, fast and scalable data transfer as well as generality in terms of heterogeneous systems and environments. Hence, in this paper, we introduce a holistic data transfer framework. Our XDBC framework splits the data transfer pipeline into logical components and provides a wide variety of physical implementations for these components. This design allows a seamless integration of different systems as well as the automatic optimizations of data transfer configurations according to workload and environment characteristics. Our evaluation shows that XDBC outperforms state-of-the-art generic data transfer tools by up to 5x, while being on par with specialized approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- Skyplane: Optimizing Transfer Cost and Throughput Using Cloud-Aware OverlaysParas Jain, Sam Kumar, Sarah Wooders, Shishir G. Patil 等NSDI 2023 · 被引用 66 次
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 被引用 45 次
- The Composable Data Management System ManifestoPedro Pedreira, Orri Erling, Konstantinos Karanasos, Scott Schneider 等VLDB 2023 · 被引用 36 次
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 被引用 32 次
相关 Paper
- In-Situ Cross-Database Query ProcessingHaralampos Gavriilidis, Kaustubh Beedkar, Jorge-Arnulfo Quiané-Ruiz, Volker MarklICDE 2023 · 被引用 12 次
- Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science EngineWeizheng Lu, Chao Hui, Yunhai Wang, Feng Zhang 等VLDB 2025 · 被引用 1 次
- ConnectorX: Accelerating Data Loading From Databases to DataframesXiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen 等VLDB 2022 · 被引用 14 次
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!William Zhang, Wan Shen Lim, Andrew PavloSIGMOD 2026 · 被引用 7 次
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 被引用 13 次
