Automated Translation of Functional Big Data Queries to SQL
Guoqiang Zhang, Benjamin Mariano, Xipeng Shen, Isil Dillig
摘要
Big data analytics frameworks like Apache Spark and Flink enable users to implement queries over large, distributed databases using functional APIs. In recent years, these APIs have grown in popularity because their functional interfaces abstract away much of the minutiae of distributed programming required by traditional query languages like SQL. However, the convenience of these APIs comes at a cost because functional queries are often less efficient than their SQL counterparts. Motivated by this observation, we present a new technique for automatically transpiling functional queries to SQL. While our approach is based on the standard paradigm of counterexample-guided inductive synthesis, it uses a novel column-wise decomposition technique to split the synthesis task into smaller subquery synthesis problems. We have implemented this approach as a new tool called RDD2SQL for translating Spark RDD queries to SQL and empirically evaluate the effectiveness of RDD2SQL on a set of real-world RDD queries. Our results show that (1) most RDD queries can be translated to SQL, (2) our tool is very effective at automating this translation, and (3) performing this translation offers significant performance benefits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SQL Engines Excel at the Execution of Imperative ProgramsTim Fischer, Denis Hirn, Torsten GrustVLDB 2024 · 被引用 3 次
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan 等SIGMOD 2025 · 被引用 2 次
- Optimal Predicate Pushdown SynthesisRobert Zhang, Eric Hayden Campbell, Dixin Tang, Isil DilligPLDI 2026 · 被引用 1 次
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang 等OOPSLA 2025
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
它引用的顶会 Paper7
- Data Migration using Datalog Program SynthesisYuepeng Wang, Rushi Shah, Abby Criswell, Rong Pan 等VLDB 2020 · 被引用 30 次
- Demystifying Loops in Smart ContractsBenjamin Mariano, Yanju Chen, Yu Feng, Shuvendu K. Lahiri 等ASE 2020 · 被引用 20 次
- PATSQL: Efficient Synthesis of SQL Queries from Example Tables with Quick Inference of Projected ColumnsKeita Takenouchi, Takashi Ishio, Joji Okada, Yuji SakataVLDB 2021 · 被引用 19 次
- UDF to SQL translation through compositional lazy inductive synthesisGuoqiang Zhang, Yuanchao Xu, Xipeng Shen, Isil DilligOOPSLA 2021 · 被引用 14 次
- Example-guided synthesis of relational queriesAalok Thakkar, Aaditya Naik, Nathaniel Sands, Rajeev Alur 等PLDI 2021 · 被引用 12 次
相关 Paper
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 被引用 13 次
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi 等OSDI 2020 · 被引用 5 次
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 被引用 11 次
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu 等VLDB 2026
- Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into SparkBowen Yu, Guanyu Feng, Huanqi Cao, Xiaohan Li 等VLDB 2022 · 被引用 3 次
