SubTokenTest: A Practical Benchmark for Real-World Sub-token Understanding
Shuyang Hou, Yi Hu, Muhan Zhang
摘要
Recent advancements in large language models (LLMs) have significantly enhanced their reasoning capabilities. However, they continue to struggle with basic character-level tasks, such as counting letters in words, a problem rooted in their tokenization process. While existing benchmarks have highlighted this weakness through basic character operations, such failures are often dismissed due to lacking practical relevance. Yet, many real-world applications, such as navigating text-based maps or interpreting structured tables, rely heavily on precise sub-token understanding. In this regard, we introduce SubTokenTest, a comprehensive benchmark that assesses sub-token understanding through practical, utility-driven tasks. Our benchmark includes ten tasks across four domains and isolates tokenization-related failures by decoupling performance from complex reasoning. We provide a comprehensive evaluation of nine advanced LLMs. Additionally, we investigate the impact of test-time scaling on sub-token reasoning and explore how character-level information is encoded within the hidden states.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen 等ACL 2025 · 被引用 116 次
- Case-Based or Rule-Based: How Do Transformers Do the Math?Yi Hu, Xiaojuan Tang, Haotong Yang, Muhan ZhangICML 2024 · 被引用 34 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- StochasTok: Improving Fine-Grained Subword Understanding in LLMsAnya Sims, Thomas Foster, T. Duy Nguyen-Hien, Klara Kaleb 等ICLR 2026 · 被引用 8 次
- Beyond Single-Task: Robust Multi-Task Length Generalization for LLMsYi Hu, Shijia Kang, Haotong Yang, Haotian Xu 等NeurIPS 2025 · 被引用 6 次
相关 Paper
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 被引用 3 次
- The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language ModelsAdrian Cosma, Stefan Ruseti, Emilian Radoi, Mihai DascaluEMNLP 2025 · 被引用 13 次
- TASE: Token Awareness and Structured Evaluation for Multilingual Language ModelsChenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu 等AAAI 2026 · 被引用 1 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- Number Cookbook: Number Understanding of Language Models and How to Improve ItHaotong Yang, Yi Hu, Shijia Kang, Zhouchen Lin 等ICLR 2025
