Efficient Algorithms for the Uniform Tokenization Problem
Angela W. Li, Konstantinos Mamouras
2025年份
3被引次数
3顶会引用
摘要
Tokenization (also known as scanning or lexing) is a computational task that has applications in the lexical analysis of programs during compilation and in data extraction and analysis for unstructured or semistructured data (e.g., data represented using the JSON and CSV data formats). We propose two algorithms for the tokenization problem that have linear time complexity (in the length of the input text) without using large amounts of memory. We also show that an optimized version of one of these algorithms performs well compared to prior approaches on practical tokenization workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Streaming Validation of JSON Documents Against SchemasAlexis Le Glaunec, Angela W. Li, Konstantinos MamourasVLDB 2026 · 被引用 2 次
- Static Analysis for Efficient Streaming TokenizationAngela W. Li, Yudi Yang, Konstantinos MamourasASPLOS 2026 · 被引用 1 次
- Formally Verified Linear-Time Invertible LexingSamuel Chassot, Viktor KuncakCAV 2026
它引用的顶会 Paper6
- Software-hardware codesign for efficient in-memory regular pattern matchingLingkun Kong, Qixuan Yu, Agnishom Chattopadhyay, Alexis Le Glaunec 等PLDI 2022 · 被引用 23 次
- Regular Expression Matching using Bit Vector AutomataAlexis Le Glaunec, Lingkun Kong, Konstantinos MamourasOOPSLA 2023 · 被引用 22 次
- Efficient Matching of Regular Expressions with Lookaround AssertionsKonstantinos Mamouras, Agnishom ChattopadhyayPOPL 2024 · 被引用 19 次
- Linear Matching of JavaScript Regular ExpressionsAurèle Barrière, Clément Pit-ClaudelPLDI 2024 · 被引用 11 次
- BVAP: Energy and Memory Efficient Automata Processing for Regular Expressions with Bounded RepetitionsZiyuan Wen, Lingkun Kong, Alexis Le Glaunec, Konstantinos Mamouras 等ASPLOS 2024 · 被引用 11 次
相关 Paper
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 被引用 6 次
- Tokenisation is NP-CompletePhilip Whittington, Gregor Bachmann, Tiago PimentelACL 2025 · 被引用 6 次
- LoPT: Lossless Parallel Tokenization Acceleration for Long Context Inference of Large Language ModelWei Shao, Lingchao Zheng, Pengyu Wang, Peizhen Zheng 等ACL 2026 · 被引用 1 次
- Tokenisation over Bounded Alphabets is HardVioleta Kastreva, Philip Whittington, Dennis Komm, Tiago PimentelICLR 2026 · 被引用 6 次
- An Efficient Algorithm for Streaming BPE TokenizationKonstantinos Mamouras, Angela W. Li, Yudi YangPLDI 2026
