ConnectorX: Accelerating Data Loading From Databases to Dataframes
Xiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen, Nick Zrymiak, Changbo Qu, Lampros Flokas, George Chow, Jiannan Wang, Tianzheng Wang, Eugene Wu, Qingqing Zhou
摘要
Data is often stored in a database management system (DBMS) but dataframe libraries are widely used among data scientists. An important but challenging problem is how to bridge the gap between databases and dataframes. To solve this problem, we present ConnectorX, a client library that enables fast and memory-efficient data loading from various databases to different dataframes. We first investigate why the loading process is slow and consumes large memory. We surprisingly find that the main overhead comes from the client-side rather than query execution or data transfer. We integrate several existing and new techniques to reduce the overhead and carefully design the system architecture and interface to make ConnectorX easy to extend to various databases and dataframes. Moreover, we propose server-side result partitioning that can be adopted by DBMSs in order to better support exporting data to data science tools. We conduct extensive experiments to evaluate ConnectorX and compare it with popular libraries. The results show that ConnectorX significantly outperforms existing solutions. ConnectorX is open sourced at: https://github.com/sfu-db/connector-x.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Fast and Scalable Data Transfer Across Data SystemsHaralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker MarklSIGMOD 2025 · 被引用 4 次
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu 等SIGMOD 2026 · 被引用 2 次
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm 等VLDB 2026
它引用的顶会 Paper3
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File FormatsTianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim 等VLDB 2021 · 被引用 28 次
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm 等SIGMOD 2020 · 被引用 27 次
相关 Paper
- SplitDF: Splitting Dataframes for Memory-Efficient Data AnalysisAarati Kakaraparthy, Jignesh M. PatelVLDB 2024
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 被引用 4 次
- Dias: Dynamic Rewriting of Pandas CodeStefanos Baziotis, Daniel D. Kang, Charith MendisSIGMOD 2024 · 被引用 7 次
- Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe SystemDevin Petersohn, Dixin Tang, Rehan Sohail Durrani, Areg Melik-Adamyan 等VLDB 2022 · 被引用 22 次
- PolyFrame: A Retargetable Query-based Approach to Scaling DataframesPhanwadee Sinthong, Michael J. CareyVLDB 2021 · 被引用 8 次
