Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data
Shihui Xu, Jiayi Wang, Guoliang Li
Abstract
Large language models (LLMs) have revolutionized the semantic processing of unstructured data, unlocking new avenues for data analytics beyond traditional relational databases. However, optimizing query execution over unstructured data requires accurate semantic cardinality estimation, a problem that remains unsolved due to the absence of schemas and the high cost of semantic judgments. In this paper, we present SemStats, a framework that efficiently and accurately estimates the number of results for semantic queries. SemStats first constructs a semantic catalog that identifies core semantic dimensions and their relationships within the corpus. Based on the catalog, SemStats then builds a semantic index that constructs the relationships between data points and semantic catalog. As it is expensive to assess the relationships between data points and catalog using LLMs, we propose specialized lightweight models to improve index-construction cost. For online estimation, SemStats employs stratified importance sampling guided by the semantic index and a lightweight judge model to refine query-specific evaluations. Experiments on multiple real-world datasets show that SemStats substantially reduces estimation error (up to 25×) and latency (up to 31×) compared to state-of-the-art baselines, enabling accurate and efficient cardinality estimation of semantic queries on unstructured data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan et al.VLDB 2024 · 165 citations
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang et al.VLDB 2021 · 138 citations
Related papers
- UQE: A Query Engine for Unstructured DatabasesHanjun Dai, Bethany Wang, Xingchen Wan, Bo Dai et al.NeurIPS 2024 · 45 citations
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang et al.VLDB 2026 · 5 citations
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query OptimizationJunhao Zhu, Lu Chen, Xiangyu Ke, Ziquan Fang et al.SIGMOD 2026 · 8 citations
- Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic ExplorationQiyao Sun, Xingming Li, Xixiang He, Ao Cheng et al.AAAI 2026 · 1 citation
- Text-to-ES Bench: A Comprehensive Benchmark for Converting Natural Language to Elasticsearch QueryDongge Xue, Zhili Pu, Zhentao Xia, Hongli Sun et al.ACL 2025
