ACL2026
Multilingual Tokenization through the Lens of Indian Languages: Challenges and Insights
Maharaj Brahma, N. J. Karthika, Rajat Verma, Nagasai Saketh Naidu, Rohit Saluja, Maunendra Sankar Desarkar, Ganesh Ramakrishnan
3 citations
Abstract
Tokenization plays a pivotal role in NLP and is fundamental to training language models. How ever, existing tokenizers are often skewed to wards highresource languages, limiting their effectiveness for linguistically diverse and mor phologically rich languages such as those in the Indian subcontinent. In this work, we present a comprehensive empirical study of multilingual tokenization across 17 Indic languages span ning 11 scripts and two language families. We systematically evaluate the effects of (i) widely used subword algorithms: BPE (Sennrich et al., 2016) and Unigram LM (Kudo, 2018), (ii) script and orthographyaware normalization, (iii) vo cabulary size, and (iv) multilingual vocabu lary construction strategies. We use a combi nation of intrinsic and extrinsic evaluations to obtain the following observations: (i) script specific normalization improves tokenization quality, (ii) Unigram LM better preserves mor phological boundaries than BPE, (iii) cluster based vocabulary construction shows improve ment in downstream tasks compared to the joint method. Our findings highlight the importance of linguistically informed design choices in multilingual tokenization and offer practical guidance for building effective tokenizers for lowresource and morphologically complex lan guages.