Dias: Dynamic Rewriting of Pandas Code
Stefanos Baziotis, Daniel D. Kang, Charith Mendis
摘要
In recent years, dataframe libraries, such as pandas have exploded in popularity. Due to their flexibility, they are increasingly used in ad-hoc exploratory data analysis (EDA) workloads. These workloads are diverse, including custom functions which can span libraries or be written in pure Python. The majority of systems available to accelerate EDA workloads focus on bulk-parallel workloads, which contain vastly different computational patterns, typically within a single library. As a result, they can introduce excessive overheads for ad-hoc EDA workloads due to their expensive optimization techniques. Instead, we identify source-to-source, external program rewriting as a lightweight technique which can optimize across representations, and offer substantial speedups while also avoiding slowdowns. We implemented Dias, which rewrites notebook cells to be more efficient for ad-hoc EDA workloads. We develop techniques for efficient rewrites in Dias, including checking the preconditions under which rewrites are correct, dynamically, at fine-grained program points. We show that Dias can rewrite individual cells to be 57× faster compared to pandas and 1909× faster compared to optimized systems such as modin. Furthermore, Dias can accelerate whole notebooks by up to 3.6× compared to pandas and 27.1× compared to modin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Optimal Predicate Pushdown SynthesisRobert Zhang, Eric Hayden Campbell, Dixin Tang, Isil DilligPLDI 2026 · 被引用 1 次
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang 等OOPSLA 2025
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
它引用的顶会 Paper7
- Alive2: bounded translation validation for LLVMNuno P. Lopes, Juneyoung Lee, Chung-Kil Hur, Zhengyang Liu 等PLDI 2021 · 被引用 109 次
- PET: Optimizing Tensor Programs with Partially Equivalent Transformations and Automated CorrectionsHaojie Wang, Jidong Zhai, Mingyu Gao, Zixuan Ma 等OSDI 2021 · 被引用 77 次
- Lux: Always-on Visualization Recommendations for Exploratory Dataframe WorkflowsDoris Jung Lin Lee, Dixin Tang, Kunal Agarwal, Thyne Boonmark 等VLDB 2022 · 被引用 61 次
- VeGen: a vectorizer generator for SIMD and beyondYishen Chen, Charith Mendis, Michael Carbin, Saman P. AmarasingheASPLOS 2021 · 被引用 43 次
- Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe SystemDevin Petersohn, Dixin Tang, Rehan Sohail Durrani, Areg Melik-Adamyan 等VLDB 2022 · 被引用 22 次
相关 Paper
- DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in PythonJinglin Peng, Weiyuan Wu, Brandon Lockhart, Song Bian 等SIGMOD 2021 · 被引用 29 次
- Towards Scalable Dataframe SystemsDevin Petersohn, William W. Ma, Doris Jung Lin Lee, Stephen Macke 等VLDB 2020 · 被引用 109 次
- PyTond: Efficient Python Data Science on the Shoulders of DatabasesHesam Shahrokhi, Amirali Kaboli, Mahdi Ghorbani, Amir ShaikhhaICDE 2024 · 被引用 4 次
- HARP: holistic analysis for refactoring Python-based analytics programsWeijie Zhou, Yue Zhao, Guoqiang Zhang, Xipeng ShenICSE 2020 · 被引用 10 次
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
