Locally Consistent Parsing for Text Indexing in Small Space
Or Birenzwige, Shay Golan, Ely Porat
Abstract
We consider two closely related problems of text indexing in a sub-linear working space. The first problem is the Sparse Suffix Tree (SST) construction, where a text S is given in read-only memory, along with a set of suffixes B, and the goal is to construct the compressed trie of all these suffixes ordered lexicographically, using only O(|B|) words of space. The second problem is the Longest Common Extension (LCE) problem, where again a text S of length n is given in read-only memory with some parameter 1 ≤ τ ≤ n, and the goal is to construct a data structure that uses O( n τ ) words of space and can compute for any pair of suffixes their longest common prefix length. We show how to use ideas based on the Locally Consistent Parsing technique, that were introduced by Sahinalp and Vishkin [44], in some non-trivial ways in order to improve the known results for the above problems. We introduce new Las-Vegas and deterministic algorithms for both problems.
For the randomized algorithms, we introduce the first Las-Vegas SST construction algorithm that takes O(n) time. This is an improvement over the last result of Gawrychowski and Kociumaka [22] who obtained O(n) time for Monte Carlo algorithm, and O(n log |B|) time with high probability for Las-Vegas algorithm. In addition, we introduce a randomized Las-Vegas construction for a data structure that uses O( n τ ) words of space, can be constructed in linear time with high probability and answers LCE queries in O(τ ) time.
For the deterministic algorithms, we introduce an SST construction algorithm that takes O(n log n |B| ) time (for |B| = Ω(log n)). This is the first almost linear time, O(n•polylog n), deterministic SST construction algorithm, where all previous algorithms take at least Ω minn|B|, n 2 |B| time. For the LCE problem, we introduce a data structure that uses O( n τ ) words of space and answers LCE queries in O(τ log * n) time, with O(n log τ ) construction time (for τ = O( n log n )). This data structure improves both query time and construction time upon the results of Tanimura et al. [47].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 970e58c6-a7a8-43b6-b6e4-8613189d5b68Cited by top-tier papers5
- Near-Optimal Quantum Algorithms for String ProblemsShyan Akmal, Ce JinSODA 2022 · 15 citations
- Text Indexing for Long Patterns: Anchors are All you NeedLorraine A. K. Ayad, Grigorios Loukides, Solon P. PissisVLDB 2023 · 10 citations
- Locally Consistent Decomposition of Strings with Applications to Edit Distance SketchingSudatta Bhattacharya, Michal KouckýSTOC 2023 · 3 citations
- Space-Efficient Text Indexing with Mismatches using Function InversionJackson Bibbens, Levi Borevitz, Samuel McCauleySTOC 2026 · 2 citations
- Dynamic suffix array with polylogarithmic queries and updatesDominik Kempa, Tomasz KociumakaSTOC 2022
Related papers
- Breaking the 𝒪(n)-Barrier in the Construction of Compressed Suffix Arrays and Suffix TreesDominik Kempa, Tomasz KociumakaSODA 2023 · 10 citations
- Collapsing the Hierarchy of Compressed Data Structures: Suffix Arrays in Optimal Compressed SpaceDominik Kempa, Tomasz KociumakaFOCS 2023 · 20 citations
- Tight Lower Bounds for Central String Queries in Compressed SpaceDominik Kempa, Tomasz KociumakaSODA 2026
- Lempel-Ziv (LZ77) Factorization in Sublinear TimeDominik Kempa, Tomasz KociumakaFOCS 2024 · 2 citations
- An Upper Bound and Linear-Space Queries on the LZ-End ParsingDominik Kempa, Barna SahaSODA 2022 · 12 citations
