Lune

SIGMOD2026顶会

EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines

Qingfeng Pan, Qingyuan Jing, Chenyang Zhang, Jiahe Zhi, Chen Xu, Heng Long

2026年份

摘要

Invoking trained machine learning (ML) pipelines to perform intelligent analysis on data stored in databases has gradually become an essential requirement for many applications. Considering performance and privacy demands, prior studies translate ML pipelines into pure SQL queries for native execution. They typically rely on translating encoders into CASE expressions. Unfortunately, databases generally do not deeply optimize CASE expressions, as they are not first-class citizens in query optimizers, which may cause inefficient performance of generated queries. To overcome this limitation, we introduce join-based translation approaches, which convert encoders into joins, thus leveraging the high-performance join execution of modern databases. Furthermore, multiple encoders may have various join translation combinations, and join-based translations do not always outperform the case-based ones. This leads to a huge space of candidate translation plans that generate queries having drastically varying execution time. Hence, we propose a plan selector to address this challenge. It employs dynamic programming-based and priority-based selection strategies to identify an efficient translation plan from the huge search space. Moreover, we implement an ML2SQL framework, namely EncoderForge, deployable as a plugin across various databases. Experimental results demonstrate that queries generated by EncoderForge achieve up to an order-of-magnitude speedup compared with those produced by existing translation approaches.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖