Comparative Analysis of the Intrinsic Metrics for Tokenizers and their effect on Downstream Tasks for Hindi and Marathi
Shagun Dwivedi, Kaushik Gopalan
摘要
Various studies have pointed out that the performance of language models is poor in non-English or non-European languages. One of the factors affecting this performance is the effectiveness and suitability of the tokenization scheme used in the model. Indic scripts require multiple Unicode codepoints to represent a single visual unit to be encoded in the standard UTF-8 scheme. This paper investigates the effect of multiple tokenizers that use UTF-8 text input on the downstream performance of pretrained language models for Hindi and Marathi, languages written in Devanāgari script. We present the intrinsic performance of the tokenizers using Fertility, Rényi Efficiency and Percentile Frequency, and report the extrinsic performance of monolingual and multilingual models on question-answering tasks, using an automated parts-of-speech and sentence similarity based evaluation framework, and on word-level tasks such as grapheme-to-phoneme conversion and transliteration. We propose a grapheme cluster tokenizer for the script which shows performance better than or competitive with other popular tokenizers. We also find that the Rényi Efficiency metric is highly correlated to downstream performance on question answering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng 等EMNLP 2023 · 被引用 91 次
- Don't Trust ChatGPT when your Question is not in English: A Study of Multilingual Abilities and Types of LLMsXiang Zhang, Senyu Li, Bradley Hauer, Ning Shi 等EMNLP 2023 · 被引用 58 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
相关 Paper
- MUTANT: A Recipe for Multilingual Tokenizer DesignSouvik Rana, Ashish Kulkarni, Arul Menezes, Chandra Khatri 等ACL 2026 · 被引用 2 次
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 被引用 4 次
- Beyond Text Compression: Evaluating Tokenizers Across ScalesJonas F. Lotz, António Vilarinho Lopes, Stephan Peitz, Hendra Setiawan 等ACL 2025 · 被引用 3 次
- Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic AlphabetMilan Miletic, Julie Kallini, Ekaterina ShutovaACL 2026
- Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan LanguagesTejas I. Dhamecha, V. Rudra Murthy, Samarth Bharadwaj, Karthik Sankaranarayanan 等EMNLP 2021 · 被引用 18 次
