Corpus-Dependent Subcharacter Encoding via HMM-Guided Code Assignment
Tatsuya Hiraoka
摘要
We propose a corpus-dependent alternative to byte encoding that learns fixed-length atomic codes for characters directly from text, which we refer to as Latom (Learned Atom-based Encoding). We instantiate this framework by training an HMM on N -repeated character sequences to estimate "atom" posteriors, followed by a Hungarian assignment yielding a globally optimal one-to-one character-code mapping. Across 14 languages, the encodings improve intrinsic metrics, including token counts after subword tokenization and bigram perplexity, with appropriate code lengths. On Amazon Reviews in six languages, Latom improves text classification accuracy and reduces decoding errors in language model generation. Overall, these results demonstrate that character encodings can be learned from corpus statistics while remaining reversible and compatible with standard tokenization pipelines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
- Dynamic Programming Encoding for Subword Segmentation in Neural Machine TranslationXuanli He, Gholamreza Haffari, Mohammad NorouziACL 2020 · 被引用 33 次
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia 等ACL 2024 · 被引用 1 次
相关 Paper
- T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient EmbeddingsBjörn Deiseroth, Manuel Brack, Patrick Schramowski, Kristian Kersting 等EMNLP 2024 · 被引用 2 次
- Coding Textual Inputs Boosts the Accuracy of Neural NetworksAbdul Rafae Khan, Jia Xu, Weiwei SunEMNLP 2020 · 被引用 2 次
- Local Byte Fusion for Neural Machine TranslationMakesh Narsimhan Sreedhar, Xiangpeng Wan, Yu Cheng, Junjie HuACL 2023 · 被引用 2 次
- ByteFlow: Language Modeling through Adaptive Byte Compression without a TokenizerChunyuan Deng, Sanket Lokegaonkar, Colin Lockard, Besnik Fetahu 等ICLR 2026 · 被引用 1 次
- Language Models over Canonical Byte-Pair EncodingsTim Vieira, Tianyu Liu, Clemente Pasti, Yahya Emara 等ICML 2025
