The Case For Language Model Approximated LIKE Predicate
Yingze Li, Dong Wang, Zixuan Wang, Yingli Zhou, Yu Yan, Jian Geng, Xinyue Wang, Ziqing Zeng, Hongzhi Wang
摘要
Modern databases often use the LIKE predicate to search text data. However, when the search condition is interrupted by wildcards, the existing search structure can degrade to a worst-case complexity linear to full table scale, resulting in poor performance. Traditional methods, such as B+-trees, fail to handle wildcards at both ends efficiently. Recent advances in language models offer a promising solution. These models can decode complex LIKE patterns into a small set of candidate values, which are then verified in dataset-size-invariant time via hash table lookups, greatly improving efficiency. However, integrating LLMs into databases faces challenges such as high latency, large storage requirements, and sensitivity to data distribution drifts. To address these issues, we propose SMILE, a S mall language M odel I ntegrated L IKE E ngine that learns column-local character distributions through small but exquisite parameters. Our SMILE acts as a neural translator that converts complex LIKE patterns into their corresponding result sets. Our approach achieves asymptotic complexity improvements while preserving SQL LIKE logic. We conduct comprehensive evaluation across diverse datasets to validate the efficacy of our approach. Our compact SMILE, with a parameter size 5 orders of magnitude smaller than large language models, achieves strong LIKE decoding efficiency and quality. Specifically, SMILE obtains high recall ability while accelerating LIKE by 3 orders of magnitude compared to large language models and sequential scans, 1.8-41.6 times faster than trigram indexes, and 2 orders of magnitude faster than B+-trees. Moreover, our model demonstrates robustness against potential data and query distribution drifts.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LPLM: A Neural Language Model for Cardinality Estimation of LIKE-QueriesMehmet Aytimur, Silvan Reiner, Leonard Wörteler, Theodoros Chondrogiannis 等SIGMOD 2024 · 被引用 13 次
- TranSQL + : Serving Large Language Models with SQL on Low-Resource HardwareWenbo Sun, Qiming Guo, Wenlu Wang, Rihan HaiSIGMOD 2026 · 被引用 1 次
- RISE: Rule-Driven SQL Dialect Translation via Query ReductionXudong Xie, Yuwei Zhang, Wensheng Dou, Yu Gao 等ICSE 2026
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
- Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained DecodingSmit Jivani, Sarvam Maheshwari, Sunita SarawagiSIGMOD 2026 · 被引用 1 次
