CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
Omri Uzan, Yuval Pinter
摘要
Tasks that require character-level reasoning, such as counting or locating characters within words, remain challenging for contemporary language models. A common conjecture is that language models' reliance on subword units, rather than characters, contributes to their struggles with character-level tasks, yet recent studies offer conflicting conclusions about the role of tokenization, leaving its impact unclear. To address this gap, we introduce CharBench, a comprehensive benchmark of character-level tasks that is two orders of magnitude larger than existing alternatives. We evaluate a diverse range of leading open-weight and proprietary models on CharBench and find that it presents a significant challenge to modern LLMs, with average accuracies of 43.6% and 32.3% on some tasks. We present an in-depth analysis of how intrinsic properties of words and their segmentations into tokens correspond to model performance. For counting tasks, we find that tokenization properties are weakly correlated with correctness, while the length of the queried word and the actual character count play a more significant part. In contrast, for tasks requiring intra-word positional understanding, performance is negatively correlated with the length of the token containing the queried character, suggesting that longer tokens obscure information on character position for LLMs. We encourage future work to build on the benchmark and evaluation methodology introduced here as tools for improving model performance on these tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
- Tokenization and the Noiseless ChannelVilém Zouhar, Clara Meister, Juan Luis Gastaldi, Li Du 等ACL 2023 · 被引用 10 次
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 被引用 6 次
- StringLLM: Understanding the String Processing Capability of Large Language ModelsXilong Wang, Hao Fu, Jindong Wang, Neil Zhenqiang GongICLR 2025
相关 Paper
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language ModelsAdrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai DascaluEMNLP 2025 · 被引用 13 次
- TASE: Token Awareness and Structured Evaluation for Multilingual Language ModelsChenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu 等AAAI 2026 · 被引用 1 次
- StochasTok: Improving Fine-Grained Subword Understanding in LLMsAnya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb 等ICLR 2026 · 被引用 8 次
- C-LLM: Learn to Check Chinese Spelling Errors Character by CharacterKunting Li, Yong Hu, Liang He, Fandong Meng 等EMNLP 2024 · 被引用 9 次
