Featurized-Decomposition Join: Low-Cost Semantic Joins with Guarantees
Sepanta Zeighami, Shreya Shankar, Aditya G. Parameswaran
摘要
Large Language Models (LLMs) are being increasingly used within data systems to process large datasets with text fields. A broad class of such tasks involves a semantic join —joining two tables based on a natural language predicate per pair of tuples, evaluated using an LLM. Semantic joins generalize tasks such as entity matching and record categorization, as well as more complex text understanding tasks. A naive implementation is expensive as it requires invoking an LLM for every pair of rows in the cross product. Existing approaches mitigate this cost by first applying embedding-based semantic similarity to filter candidate pairs, deferring to an LLM only when similarity scores are deemed inconclusive. However, these methods yield limited gains in practice, since semantic similarity may not reliably predict the join outcome. We propose Featurized-Decomposition Join (FDJ for short), a novel approach for performing semantic joins that significantly reduces cost while preserving quality. FDJ automatically extracts features and combines them into a logical expression in conjunctive normal form that we call a featurized decomposition to effectively prune out non-matching pairs. A featurized decomposition extracts key information from text records and performs inexpensive comparisons on the extracted features. We show how to use LLMs to automatically extract reliable features and compose them into logical expressions while providing statistical guarantees on the output—an inherently challenging problem due to dependencies among features. Experiments show up to 10 times reduction in cost compared to the state-of-the-art while providing the same guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- On the Theoretical Limitations of Embedding-Based RetrievalOrion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk LeeICLR 2026 · 被引用 138 次
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani 等VLDB 2021 · 被引用 109 次
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 被引用 103 次
相关 Paper
- DeepJoin: Joinable Table Discovery with Pre-trained Language ModelsYuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto 等VLDB 2023 · 被引用 53 次
- QUEST: Query Optimization in Unstructured Document AnalysisZhaoze Sun, Chengliang Chai, Qiyan Deng, Kaisen Jin 等VLDB 2025 · 被引用 9 次
- Factorized and Vectorized Execution: Optimizing Analytical and Semantic Queries over RelationsSunny Yasser, Anas Dorbani, Amine MhedhbiSIGMOD 2026 · 被引用 1 次
- In-context Clustering-based Entity Resolution with Large Language Models: A Design Space ExplorationJiajie Fu, Haitong Tang, Arijit Khan, Sharad Mehrotra 等SIGMOD 2026 · 被引用 8 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
