Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System
Devin Petersohn, Dixin Tang, Rehan Sohail Durrani, Areg Melik-Adamyan, Joseph Gonzalez, Anthony D. Joseph, Aditya G. Parameswaran
Abstract
Dataframes have become universally popular as a means to represent data in various stages of structure, and manipulate it using a rich set of operators---thereby becoming an essential tool in the data scientists' toolbox. However, dataframe systems, such as pandas, scale poorly---and are non-interactive on moderate to large datasets. We discuss our experiences developing Modin, our first cut at a parallel dataframe system, which already has users across several industries and over 1M downloads. Modin translates pandas functions into a core set of operators that are individually parallelized via columnar, row-wise, or cell-wise decomposition rules that we formalize in this paper. We also introduce metadata independence to allow metadata---such as order and type---to be decoupled from the physical representation and maintained lazily. Using rule-based decomposition and metadata independence, along with careful engineering, Modin is able to support pandas operations across both rows and columns on very large dataframes---unlike Koalas and Dask DataFrames that either break down or are unable to support such operations, while also being much faster than pandas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e733a505-1b3d-4650-ab3a-db983545280bCited by top-tier papers7
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 15 citations
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 13 citations
- HiPy: Extracting High-Level Semantics from Python Code for Data ProcessingMichael Jungmair, Alexis Engelke, Jana GicevaOOPSLA 2024 · 7 citations
- Dias: Dynamic Rewriting of Pandas CodeStefanos Baziotis, Daniel D. Kang, Charith MendisSIGMOD 2024 · 7 citations
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan et al.SIGMOD 2025 · 2 citations
Builds on1
Related papers
- DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in PythonJinglin Peng, Weiyuan Wu, Brandon Lockhart, Song Bian et al.SIGMOD 2021 · 29 citations
- Moko: Marrying Python with Big Data SystemsKe Meng, Tao He, Sijie Shen, Lei Wang et al.EuroSys 2025 · 2 citations
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 4 citations
- ConnectorX: Accelerating Data Loading From Databases to DataframesXiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen et al.VLDB 2022 · 14 citations
