Lune

VLDB2026顶会

SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora

Pranay Mundra, Daniel Kocher, Martin Schaeler, Nikolaus Augsten

2026年份

摘要

A two-stage pipeline is commonly used to identify similar text passages in large document corpora: First, a fast approach produces potential matches, which are then examined in detail. Existing approaches for the first step consider only syntactic information and miss semantically similar passages that are syntactically dissimilar. To address this, we define the novel problem of semantic document alignment as a semantic set-similarity problem on k -width windows. For two documents S and T, an exhaustive baseline that evaluates all |S| × |T| window pairs is computationally infeasible since assessing the similarity of a single pair requires O ( k 3 ) time.

We propose SeDA, which combines a sophisticated candidate generation technique with a bound cascade to drastically reduce the number of expensive window comparisons. It further exploits overlapping windows to efficiently compute both the bounds and the final similarity scores. Our empirical results on three large document corpora indicate that SeDA prunes over 99% of the window similarity computations, resulting in response-time improvements of 1.5–3 orders of magnitude over the baseline solution and 2–5 orders of magnitude over SBERT. Compared to purely syntactic competitors, SeDA provides competitive runtimes and achieves superior result quality, i.e., near-optimal F1-Score of precision/recall and matching the performance of purely semantic methods such as SBERT.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 51c326d5-3f32-48c2-af59-57ab35f8173c

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖