Automated Translation of Functional Big Data Queries to SQL
Guoqiang Zhang, Benjamin Mariano, Xipeng Shen, Isil Dillig
Abstract
Big data analytics frameworks like Apache Spark and Flink enable users to implement queries over large, distributed databases using functional APIs. In recent years, these APIs have grown in popularity because their functional interfaces abstract away much of the minutiae of distributed programming required by traditional query languages like SQL. However, the convenience of these APIs comes at a cost because functional queries are often less efficient than their SQL counterparts. Motivated by this observation, we present a new technique for automatically transpiling functional queries to SQL. While our approach is based on the standard paradigm of counterexample-guided inductive synthesis, it uses a novel column-wise decomposition technique to split the synthesis task into smaller subquery synthesis problems. We have implemented this approach as a new tool called RDD2SQL for translating Spark RDD queries to SQL and empirically evaluate the effectiveness of RDD2SQL on a set of real-world RDD queries. Our results show that (1) most RDD queries can be translated to SQL, (2) our tool is very effective at automating this translation, and (3) performing this translation offers significant performance benefits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a189a652-4b03-405a-a689-0513461bfdd7Cited by top-tier papers5
- SQL Engines Excel at the Execution of Imperative ProgramsTim Fischer, Denis Hirn, Torsten GrustVLDB 2024 · 3 citations
- QURE: AI-Assisted and Automatically Verified UDF InliningTarique Siddiqui, Arnd Christian König, Jiashen Cao, Cong Yan et al.SIGMOD 2025 · 2 citations
- Optimal Predicate Pushdown SynthesisRobert Zhang, Eric Hayden Campbell, Dixin Tang, Isil DilligPLDI 2026 · 1 citation
- Homomorphism Calculus for User-Defined AggregationsZiteng Wang, Ruijie Fang, Linus Zheng, Dixin Tang et al.OOPSLA 2025
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
Builds on7
- Data Migration using Datalog Program SynthesisYuepeng Wang, Rushi Shah, Abby Criswell, Rong Pan et al.VLDB 2020 · 30 citations
- Demystifying Loops in Smart ContractsBenjamin Mariano, Yanju Chen, Yu Feng, Shuvendu K. Lahiri et al.ASE 2020 · 20 citations
- PATSQL: Efficient Synthesis of SQL Queries from Example Tables with Quick Inference of Projected ColumnsKeita Takenouchi, Takashi Ishio, Joji Okada, Yuji SakataVLDB 2021 · 19 citations
- UDF to SQL translation through compositional lazy inductive synthesisGuoqiang Zhang, Yuanchao Xu, Xipeng Shen, Isil DilligOOPSLA 2021 · 14 citations
- Example-guided synthesis of relational queriesAalok Thakkar, Aaditya Naik, Nathaniel Sands, Rajeev Alur et al.PLDI 2021 · 12 citations
Related papers
- Translation of Array-Based Loops to Distributed Data-Parallel ProgramsLeonidas Fegaras, Md Hasanuzzaman NoorVLDB 2020 · 13 citations
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi et al.OSDI 2020 · 5 citations
- Dynamic Speculative Optimizations for SQL Compilation in Apache SparkFilippo Schiavio, Daniele Bonetta, Walter BinderVLDB 2020 · 11 citations
- MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration TuningBeicheng Xu, Lingching Tung, Yuchen Wang, Yupeng Lu et al.VLDB 2026
- Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into SparkBowen Yu, Guanyu Feng, Huanqi Cao, Xiaohan Li et al.VLDB 2022 · 3 citations
