Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet
Milan Miletic, Julie Kallini, Ekaterina Shutova
摘要
Multilingual language models often exhibit performance disparities across languages that can arise as early as the tokenization stage. Widely-used subword tokenization approaches favor high-resource languages, and tokenizerfree methods still yield longer sequences for scripts with a higher bytes-per-character ratio. To address these shortcomings, we propose to use the International Phonetic Alphabet (IPA) as a language-agnostic input representation for multilingual tokenizers. IPA provides a compact symbol inventory, greater cross-lingual character overlap, and a more balanced byte-per-character distribution across languages. We train matched pairs of text vs. IPA subword tokenizers across 24 languages and 14 scripts and demonstrate that IPA tokenizers consistently improve tokenization quality, especially for non-Latin scripts, and generalize more effectively to unseen languages and scripts. github.com/Mikki99/ipa-tokenization
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan 等NeurIPS 2023 · 被引用 197 次
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesPedro Javier Ortiz Suárez, Laurent Romary, Benoît SagotACL 2020 · 被引用 72 次
相关 Paper
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder 等ACL 2021
- MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language ModelingTomasz Limisiewicz, Terra Blevins, Hila Gonen, Orevaoghene Ahia 等ACL 2024 · 被引用 1 次
- MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Valentin Hofmann 等NeurIPS 2024 · 被引用 37 次
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan 等ACL 2025 · 被引用 3 次
- Unsupervised Tokenization LearningAnton Kolonin, Vignav RameshEMNLP 2022 · 被引用 3 次
