Enhancing Large Language Models through Adaptive Tokenizers
Mengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu, Binfan Zheng, Chang Xu, Yunhe Wang
摘要
Tokenizers serve as crucial interfaces between models and linguistic data, substantially influencing the efficacy and precision of large language models (LLMs). Traditional tokenization methods often rely on static frequency-based statistics and are not inherently synchronized with LLM architectures, which may limit model performance. In this study, we propose a simple but effective method to learn tok-enizers specifically engineered for seamless integration with LLMs. Initiating with a broad initial vocabulary, we refine our tokenizer by monitoring changes in the model’s perplexity during training, allowing for the selection of a tokenizer that is closely aligned with the model’s evolving dynamics. Through iterative refinement, we develop an optimized tokenizer. Our empirical evaluations demonstrate that this adaptive approach significantly enhances accuracy compared to conventional methods, maintaining comparable vocabulary sizes and affirming its potential to improve LLM functionality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- From Bytes to Ideas: Language Modeling with Autoregressive U-NetsMathurin Videau, Badr Youbi Idrissi, Alessandro Ferreira Leite, Marc Schoenauer 等NeurIPS 2025 · 被引用 14 次
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 被引用 5 次
- Skill Neologisms: Towards Skill-based Continual LearningAntonin Berthon, Nicolás Astorga, Mihaela van der SchaarICML 2026 · 被引用 1 次
- SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language ModelsMing Chen, Wenyao Li, Chao Liang, Shi Gu 等ACL 2026
- Date Fragments: A Hidden Bottleneck of Tokenization for Temporal ReasoningGagan Bhatia, Maxime Peyrard, Wei ZhaoEMNLP 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
相关 Paper
- Exploiting Vocabulary Frequency Imbalance in Language Model Pre-trainingWoojin Chung, Jeonghoon KimNeurIPS 2025 · 被引用 6 次
- Ted-Tok: Maintaining an Evolving Vocabulary for Lifelong LearningJiameng Huang, Zhi Zhang, Zhenyu He, Jiacheng Sun 等ACL 2026
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 被引用 7 次
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard 等NeurIPS 2025 · 被引用 5 次
- Getting the most out of your tokenizer for pre-training and domain adaptationGautier Dagan, Gabriel Synnaeve, Baptiste RozièreICML 2024 · 被引用 68 次
