ConnectorX: Accelerating Data Loading From Databases to Dataframes
Xiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen, Nick Zrymiak, Changbo Qu, Lampros Flokas, George Chow, Jiannan Wang, Tianzheng Wang, Eugene Wu, Qingqing Zhou
Abstract
Data is often stored in a database management system (DBMS) but dataframe libraries are widely used among data scientists. An important but challenging problem is how to bridge the gap between databases and dataframes. To solve this problem, we present ConnectorX, a client library that enables fast and memory-efficient data loading from various databases to different dataframes. We first investigate why the loading process is slow and consumes large memory. We surprisingly find that the main overhead comes from the client-side rather than query execution or data transfer. We integrate several existing and new techniques to reduce the overhead and carefully design the system architecture and interface to make ConnectorX easy to extend to various databases and dataframes. Moreover, we propose server-side result partitioning that can be adopted by DBMSs in order to better support exporting data to data science tools. We conduct extensive experiments to evaluate ConnectorX and compare it with popular libraries. The results show that ConnectorX significantly outperforms existing solutions. ConnectorX is open sourced at: https://github.com/sfu-db/connector-x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23f63fa7-0986-449b-8dd5-d16d7d3f58aaCited by top-tier papers4
- Fast and Scalable Data Transfer Across Data SystemsHaralampos Gavriilidis, Kaustubh Beedkar, Matthias Boehm, Volker MarklSIGMOD 2025 · 4 citations
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu et al.SIGMOD 2026 · 2 citations
- Accio: Bolt-on Query FederationXiaoying Wang, Jiannan Wang, Tianzheng Wang, Yong ZhangVLDB 2025
- Matryoshka: Uncovering Relevant Features in Data Lakes to Enhance Machine Learning ApplicationsFedor Turchenko, Runjie Zhang, Binger Chen, Matthias Boehm et al.VLDB 2026
Builds on3
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke et al.VLDB 2020 · 109 citations
- Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File FormatsTianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim et al.VLDB 2021 · 28 citations
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm et al.SIGMOD 2020 · 27 citations
Related papers
- SplitDF: Splitting Dataframes for Memory-Efficient Data AnalysisAarati Kakaraparthy, Jignesh M. PatelVLDB 2024
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 4 citations
- Dias: Dynamic Rewriting of Pandas CodeStefanos Baziotis, Daniel D. Kang, Charith MendisSIGMOD 2024 · 7 citations
- Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe SystemDevin Petersohn, Dixin Tang, Rehan Sohail Durrani, Areg Melik-Adamyan et al.VLDB 2022 · 22 citations
- PolyFrame: A Retargetable Query-based Approach to Scaling DataframesPhanwadee Sinthong, Michael J. CareyVLDB 2021 · 8 citations
