Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and Beyond
Siyang Liu, Naihao Deng, Sahand Sabour, Yilin Jia, Minlie Huang, Rada Mihalcea
Abstract
We propose task-adaptive tokenization 1 as a way to adapt the generation pipeline to the specifics of a downstream task and enhance long-form generation in mental health. Inspired by insights from cognitive science, our task-adaptive tokenizer samples variable segmentations from multiple outcomes, with sampling probabilities optimized based on taskspecific data. We introduce a strategy for building a specialized vocabulary and introduce a vocabulary merging protocol that allows for the integration of task-specific tokens into the pre-trained model's tokenization step. Through extensive experiments on psychological question-answering tasks in both Chinese and English, we find that our task-adaptive tokenization approach brings a significant improvement in generation performance while using up to 60% fewer tokens. Preliminary experiments point to promising results when using our tokenization approach with very large language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6c1124f-151c-4204-a715-8017abaaff99Cited by top-tier papers5
- Enhancing Large Language Models through Adaptive TokenizersMengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu et al.NeurIPS 2024 · 11 citations
- zip2zip: Inference-Time Adaptive Tokenization via Online CompressionSaibo Geng, Nathan Ranchin, Yunzhen Yao, Maxime Peyrard et al.NeurIPS 2025 · 5 citations
- Knowledge Planning in Large Language Models for Domain-Aligned Counseling SummarizationAseem Srivastava, Smriti Joshi, Tanmoy Chakraborty, Md. Shad AkhtarEMNLP 2024 · 3 citations
- Responsible Evaluation of AI for Mental HealthHiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt et al.ACL 2026
- Ted-Tok: Maintaining an Evolving Vocabulary for Lifelong LearningJiameng Huang, Zhi Zhang, Zhenyu He, Jiacheng Sun et al.ACL 2026
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TurboTransformers: an efficient GPU serving system for transformer modelsJiarui Fang, Yang Yu, Chengduo Zhao, Jie ZhouPPoPP 2021 · 117 citations
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale PretrainingWeizhen Qi, Yeyun Gong, Jian Jiao, Yu Yan et al.ICML 2021 · 54 citations
Related papers
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask GenerationLingfeng Wang, Hualing Lin, Senda Chen, Tao Wang et al.NeurIPS 2025 · 5 citations
- Learning Faster with Better Tokens: Parameter-Efficient Vocabulary Adaptation for Specialized Text SummarizationGunjan Balde, Soumyadeep Roy, Mainack Mondal, Niloy GangulyACL 2026
- End-to-End Vision Tokenizer TuningWenxuan Wang, Fan Zhang, Yufeng Cui, Haiwen Diao et al.NeurIPS 2025 · 6 citations
- Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language ModelsBoheng Sheng, Jiacheng Yao, Meicong Zhang, Guoxiu HeACL 2025 · 9 citations
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
