Bridging the Gap: Cardinality Estimation for Semantic Queries on Unstructured Data
Shihui Xu, Jiayi Wang, Guoliang Li
摘要
Large language models (LLMs) have revolutionized the semantic processing of unstructured data, unlocking new avenues for data analytics beyond traditional relational databases. However, optimizing query execution over unstructured data requires accurate semantic cardinality estimation, a problem that remains unsolved due to the absence of schemas and the high cost of semantic judgments. In this paper, we present SemStats, a framework that efficiently and accurately estimates the number of results for semantic queries. SemStats first constructs a semantic catalog that identifies core semantic dimensions and their relationships within the corpus. Based on the catalog, SemStats then builds a semantic index that constructs the relationships between data points and semantic catalog. As it is expensive to assess the relationships between data points and catalog using LLMs, we propose specialized lightweight models to improve index-construction cost. For online estimation, SemStats employs stratified importance sampling guided by the semantic index and a lightweight judge model to refine query-specific evaluations. Experiments on multiple real-world datasets show that SemStats substantially reduces estimation error (up to 25×) and latency (up to 31×) compared to state-of-the-art baselines, enabling accurate and efficient cardinality estimation of semantic queries on unstructured data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 被引用 251 次
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu 等VLDB 2020 · 被引用 206 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina 等VLDB 2020 · 被引用 154 次
- NeuroCard: One Cardinality Estimator for All TablesZongheng Yang, Amog Kamsetty, Sifei Luan, Eric Liang 等VLDB 2021 · 被引用 138 次
相关 Paper
- UQE: A Query Engine for Unstructured DatabasesHanjun Dai, Bethany Wang, Xingchen Wan, Bo Dai 等NeurIPS 2024 · 被引用 45 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query OptimizationJunhao Zhu, Lu Chen, Xiangyu Ke, Ziquan Fang 等SIGMOD 2026 · 被引用 8 次
- Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic ExplorationQiyao Sun, Xingming Li, Xixiang He, Ao Cheng 等AAAI 2026 · 被引用 1 次
- Text-to-ES Bench: A Comprehensive Benchmark for Converting Natural Language to Elasticsearch QueryDongge Xue, Zhili Pu, Zhentao Xia, Hongli Sun 等ACL 2025
