Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
Zhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang, Li Li, David Lo
Abstract
Large language models (LLMs) have shown exceptional performance in code generation and understanding tasks, yet their high computational costs hinder broader adoption. One important factor is the inherent verbosity of programming languages, such as unnecessary formatting elements and lengthy boilerplate code. This leads to inflated token counts in both input and generated outputs, which increases inference costs and slows down the generation process. Prior work improves this through simplifying programming language grammar, reducing token usage across both code understanding and generation tasks. However, it is confined to syntactic transformations, leaving significant opportunities for token reduction unrealized at the semantic level.In this work, we propose Token Sugar, a concept that replaces frequent and verbose code patterns with reversible, token-efficient shorthand in the source code. To realize this concept in practice, we designed a systematic solution that mines high-frequency, token-heavy patterns from a code corpus, maps each to a unique shorthand, and integrates them into LLM pretraining via code transformation. With this solution, we obtain 799 (code pattern, shorthand) pairs, which can reduce up to 15.1% token count in the source code and is complementary to existing syntax-focused methods. We further trained three widely used LLMs on Token Sugar-augmented data. Experimental results show that these models not only achieve significant token savings (up to 11.2% reduction) during generation but also maintain near-identical Pass@1 scores compared to baselines trained on unprocessed code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d471dff-d7ed-43a5-bcfc-cb602a1ca771Cited by top-tier papers2
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen et al.ISSTA 2026 · 1 citation
- EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code GenerationJingyu Xiao, Zhongyi Zhang, Yuxuan Wan, Yintong Huo et al.FSE 2026
Builds on15
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Diet code is healthy: simplifying programs for pre-trained models of codeZhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong GuFSE 2022 · 39 citations
- babble: Learning Better Abstractions with E-Graphs and Anti-unificationDavid Cao, Rose Kunkel, Chandrakana Nandi, Max Willsey et al.POPL 2023 · 38 citations
- Understanding neural code intelligence through program simplificationMd. Rafiqul Islam Rabin, Vincent J. Hellendoorn, Mohammad Amin AlipourFSE 2021 · 36 citations
- Probing model signal-awareness via prediction-preserving input minimizationSahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim Alain Laredo et al.FSE 2021 · 29 citations
Related papers
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM BudgetDangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo et al.ICSE 2026
- Getting the most out of your tokenizer for pre-training and domain adaptationGautier Dagan, Gabriel Synnaeve, Baptiste RozièreICML 2024 · 68 citations
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li et al.ISSTA 2024 · 8 citations
- AST-T5: Structure-Aware Pretraining for Code Generation and UnderstandingLinyuan Gong, Mostafa Elhoushi, Alvin CheungICML 2024 · 42 citations
- Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language ModelsYan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang et al.FSE 2024 · 6 citations
