SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora
Pranay Mundra, Daniel Kocher, Martin Schaeler, Nikolaus Augsten
摘要
A two-stage pipeline is commonly used to identify similar text passages in large document corpora: First, a fast approach produces potential matches, which are then examined in detail. Existing approaches for the first step consider only syntactic information and miss semantically similar passages that are syntactically dissimilar. To address this, we define the novel problem of semantic document alignment as a semantic set-similarity problem on k -width windows. For two documents S and T, an exhaustive baseline that evaluates all |S| × |T| window pairs is computationally infeasible since assessing the similarity of a single pair requires O ( k 3 ) time.
We propose SeDA, which combines a sophisticated candidate generation technique with a bound cascade to drastically reduce the number of expensive window comparisons. It further exploits overlapping windows to efficiently compute both the bounds and the final similarity scores. Our empirical results on three large document corpora indicate that SeDA prunes over 99% of the window similarity computations, resulting in response-time improvements of 1.5–3 orders of magnitude over the baseline solution and 2–5 orders of magnitude over SBERT. Compared to purely syntactic competitors, SeDA provides competitive runtimes and achieves superior result quality, i.e., near-optimal F1-Score of precision/recall and matching the performance of purely semantic methods such as SBERT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism DetectionZhizhi Wang, Chaoji Zuo, Dong DengSIGMOD 2022 · 被引用 14 次
- Allign: Aligning All-Pair Near-Duplicate Passages in Long TextsWeiqi Feng, Dong DengSIGMOD 2021 · 被引用 13 次
- Near-Duplicate Text Alignment with One Permutation HashingZhencan Peng, Yuheng Zhang, Dong DengSIGMOD 2025 · 被引用 5 次
- MetricJoin: Leveraging Metric Properties for Robust Exact Set Similarity JoinsManuel Widmoser, Daniel Kocher, Nikolaus Augsten, Willi MannICDE 2023 · 被引用 4 次
- A Two-Level Signature Scheme for Stable Set Similarity JoinsDaniel Ulrich Schmitt, Daniel Kocher, Nikolaus Augsten, Willi Mann 等VLDB 2023 · 被引用 3 次
相关 Paper
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim 等EMNLP 2020 · 被引用 126 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Scalable Attentive Sentence Pair Modeling via Distilled Sentence EmbeddingOren Barkan, Noam Razin, Itzik Malkiel, Ori Katz 等AAAI 2020 · 被引用 37 次
- Intra-Document Cascading: Learning to Select Passages for Neural Document RankingSebastian Hofstätter, Bhaskar Mitra, Hamed Zamani, Nick Craswell 等SIGIR 2021 · 被引用 35 次
- Advancing Semantic Textual Similarity Modeling: A Regression Framework with Translated ReLU and Smooth K2 LossBowen Zhang, Chunping LiEMNLP 2024 · 被引用 1 次
