PolyFrame: A Retargetable Query-based Approach to Scaling Dataframes
Phanwadee Sinthong, Michael J. Carey
Abstract
In the last few years, the field of data science has been growing rapidly as various businesses have adopted statistical and machine learning techniques to empower their decision making and applications. Scaling data analysis, possibly including the application of custom machine learning models, to large volumes of data requires the utilization of distributed frameworks. This can lead to serious technical challenges for data analysts and reduce their productivity. AFrame, a Python data analytics library, is implemented as a layer on top of Apache AsterixDB, addressing these issues by incorporating the data scientists' development environment and transparently scaling out the evaluation of analytical operations through a Big Data management system. While AFrame is able to leverage data management facilities (e.g., indexes and query optimization) and allows users to interact with a very large volume of data, the initial version only generated SQL++ queries and only operated against Apache AsterixDB. In this work, we describe a new design that retargets AFrame's incremental query formation to other query-based database systems as well, making it more flexible for deployment against other data management systems with composable query languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
- Dias: Dynamic Rewriting of Pandas CodeStefanos Baziotis, Daniel D. Kang, Charith MendisSIGMOD 2024 · 7 citations
- Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science EngineWeizheng Lu, Chao Hui, Yunhai Wang, Feng Zhang et al.VLDB 2025 · 1 citation
- SplitDF: Splitting Dataframes for Memory-Efficient Data AnalysisAarati Kakaraparthy, Jignesh M. PatelVLDB 2024
Builds on1
Related papers
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 4 citations
- Graphix: "One User's JSON is Another User's Graph"Glenn Galvizo, Michael J. CareyICDE 2024 · 1 citation
- YeSQL: "You extend SQL" with Rich and Highly Performant User-Defined Functions in Relational DatabasesYannis E. Foufoulas, Alkis Simitsis, Eleftherios Stamatogiannakis, Yannis E. IoannidisVLDB 2022 · 25 citations
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang et al.EuroSys 2025 · 2 citations
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 13 citations
