Lune

SIGMOD2026Top-tier venue

Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates

Mingxi Liu, Zhengyuan Ding, Chenyang Zhang, Qingfeng Pan, Huayou Su, Zhao Zhang, Chen Xu, Qingsong Ruan

2026Year

Abstract

In-database prediction queries that apply machine learning (ML) pipelines to perform data analysis are prevalent in many applications. Since data stored in databases is typically tabular, tree-based models are particularly well-suited and thus widely adopted for such tasks. When ML inference with a decision tree or random forest appears on a SQL predicate, existing works first perform ML inference and then determine whether the inference result satisfies the predicate. However, this leads to redundant feature tests on the tree node during the predicate evaluation. To determine whether one data record satisfies the predicate, it is possible to perform feature tests only on partial internal nodes of the tree. We identify that these redundant feature tests are caused by specific sibling and ancestor nodes. In particular, we propose the sibling-centric elimination with the merging-based subtree collapse method, and the ancestor-centric elimination with the sliding-based subtree recombination method. We implement a prototype system, called ReTree, based on DuckDB. Our experiments show that ReTree achieves a 2.56x speedup on average over DuckDB for prediction query execution and outperforms other solutions.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ba454a90-1717-42de-86a9-8f0c5cd58690

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines