Locally Consistent Parsing for Text Indexing in Small Space
Or Birenzwige, Shay Golan, Ely Porat
摘要
We consider two closely related problems of text indexing in a sub-linear working space. The first problem is the Sparse Suffix Tree (SST) construction, where a text S is given in read-only memory, along with a set of suffixes B, and the goal is to construct the compressed trie of all these suffixes ordered lexicographically, using only O(|B|) words of space. The second problem is the Longest Common Extension (LCE) problem, where again a text S of length n is given in read-only memory with some parameter 1 ≤ τ ≤ n, and the goal is to construct a data structure that uses O( n τ ) words of space and can compute for any pair of suffixes their longest common prefix length. We show how to use ideas based on the Locally Consistent Parsing technique, that were introduced by Sahinalp and Vishkin [44], in some non-trivial ways in order to improve the known results for the above problems. We introduce new Las-Vegas and deterministic algorithms for both problems.
For the randomized algorithms, we introduce the first Las-Vegas SST construction algorithm that takes O(n) time. This is an improvement over the last result of Gawrychowski and Kociumaka [22] who obtained O(n) time for Monte Carlo algorithm, and O(n log |B|) time with high probability for Las-Vegas algorithm. In addition, we introduce a randomized Las-Vegas construction for a data structure that uses O( n τ ) words of space, can be constructed in linear time with high probability and answers LCE queries in O(τ ) time.
For the deterministic algorithms, we introduce an SST construction algorithm that takes O(n log n |B| ) time (for |B| = Ω(log n)). This is the first almost linear time, O(n•polylog n), deterministic SST construction algorithm, where all previous algorithms take at least Ω minn|B|, n 2 |B| time. For the LCE problem, we introduce a data structure that uses O( n τ ) words of space and answers LCE queries in O(τ log * n) time, with O(n log τ ) construction time (for τ = O( n log n )). This data structure improves both query time and construction time upon the results of Tanimura et al. [47].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Near-Optimal Quantum Algorithms for String ProblemsShyan Akmal, Ce JinSODA 2022 · 被引用 15 次
- Text Indexing for Long Patterns: Anchors are All you NeedLorraine A. K. Ayad, Grigorios Loukides, Solon P. PissisVLDB 2023 · 被引用 10 次
- Locally Consistent Decomposition of Strings with Applications to Edit Distance SketchingSudatta Bhattacharya, Michal KouckýSTOC 2023 · 被引用 3 次
- Space-Efficient Text Indexing with Mismatches using Function InversionJackson Bibbens, Levi Borevitz, Samuel McCauleySTOC 2026 · 被引用 2 次
- Dynamic suffix array with polylogarithmic queries and updatesDominik Kempa, Tomasz KociumakaSTOC 2022
相关 Paper
- Breaking the 𝒪(n)-Barrier in the Construction of Compressed Suffix Arrays and Suffix TreesDominik Kempa, Tomasz KociumakaSODA 2023 · 被引用 10 次
- Collapsing the Hierarchy of Compressed Data Structures: Suffix Arrays in Optimal Compressed SpaceDominik Kempa, Tomasz KociumakaFOCS 2023 · 被引用 20 次
- Tight Lower Bounds for Central String Queries in Compressed SpaceDominik Kempa, Tomasz KociumakaSODA 2026
- Lempel-Ziv (LZ77) Factorization in Sublinear TimeDominik Kempa, Tomasz KociumakaFOCS 2024 · 被引用 2 次
- An Upper Bound and Linear-Space Queries on the LZ-End ParsingDominik Kempa, Barna SahaSODA 2022 · 被引用 12 次
