TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
Chenzhuo Zhao, Xinda Wang, Yue Huang, Junting Lu, Ziqian Liu
摘要
While large language models (LLMs) have demonstrated remarkable performance on high-level semantic tasks, they often struggle with fine-grained, token-level understanding and structural reasoning-capabilities that are essential for applications requiring precision and control. We introduce TASE, a comprehensive benchmark designed to evaluate LLMs' ability to perceive and reason about token-level information across languages. TASE covers 10 tasks under two core categories: token awareness and structural understanding, spanning Chinese, English, and Korean, with a 35,927-instance evaluation set and a scalable synthetic data generation pipeline for training. Tasks include character counting, token alignment, syntactic structure parsing, and length constraint satisfaction. We evaluate over 30 leading commercial and open-source LLMs, including O3, Claude 4, Gemini 2.5 Pro, and DeepSeek-R1, and train a custom Qwen2.5-14B model using the GRPO training method. Results show that human performance significantly outpaces current LLMs, revealing persistent weaknesses in token-level reasoning. TASE sheds light on these limitations and provides a new diagnostic lens for future improvements in low-level language understanding and cross-lingual generalization.Our code and dataset are publicly available at https://github.com/cyzcz/Tase .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 被引用 133 次
- MERA: A Comprehensive LLM Evaluation in RussianAlena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova 等ACL 2024 · 被引用 10 次
相关 Paper
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal VerificationXu Xu, Xin Li, Xingwei Qu, Jie Fu 等ICLR 2026 · 被引用 9 次
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei 等EMNLP 2025 · 被引用 2 次
- SubTokenTest: A Practical Benchmark for Real-World Sub-token UnderstandingShuyang Hou, Yi Hu, Muhan ZhangACL 2026
- CharBench: Evaluating the Role of Tokenization in Character-Level TasksOmri Uzan, Yuval PinterAAAI 2026 · 被引用 3 次
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning 等ICML 2025
