Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient Shorthand
Zhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang, Li Li, David Lo
摘要
Large language models (LLMs) have shown exceptional performance in code generation and understanding tasks, yet their high computational costs hinder broader adoption. One important factor is the inherent verbosity of programming languages, such as unnecessary formatting elements and lengthy boilerplate code. This leads to inflated token counts in both input and generated outputs, which increases inference costs and slows down the generation process. Prior work improves this through simplifying programming language grammar, reducing token usage across both code understanding and generation tasks. However, it is confined to syntactic transformations, leaving significant opportunities for token reduction unrealized at the semantic level.In this work, we propose Token Sugar, a concept that replaces frequent and verbose code patterns with reversible, token-efficient shorthand in the source code. To realize this concept in practice, we designed a systematic solution that mines high-frequency, token-heavy patterns from a code corpus, maps each to a unique shorthand, and integrates them into LLM pretraining via code transformation. With this solution, we obtain 799 (code pattern, shorthand) pairs, which can reduce up to 15.1% token count in the source code and is complementary to existing syntax-focused methods. We further trained three widely used LLMs on Token Sugar-augmented data. Experimental results show that these models not only achieve significant token savings (up to 11.2% reduction) during generation but also maintain near-identical Pass@1 scores compared to baselines trained on unprocessed code.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen 等ISSTA 2026 · 被引用 1 次
- EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code GenerationJingyu Xiao, Zhongyi Zhang, Yuxuan Wan, Yintong Huo 等FSE 2026
它引用的顶会 Paper15
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Diet code is healthy: simplifying programs for pre-trained models of codeZhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong GuFSE 2022 · 被引用 39 次
- babble: Learning Better Abstractions with E-Graphs and Anti-unificationDavid Cao, Rose Kunkel, Chandrakana Nandi, Max Willsey 等POPL 2023 · 被引用 38 次
- Understanding neural code intelligence through program simplificationMd. Rafiqul Islam Rabin, Vincent J. Hellendoorn, Mohammad Amin AlipourFSE 2021 · 被引用 36 次
- Probing model signal-awareness via prediction-preserving input minimizationSahil Suneja, Yunhui Zheng, Yufan Zhuang, Jim Alain Laredo 等FSE 2021 · 被引用 29 次
相关 Paper
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM BudgetDangfeng Pan, Zhensu Sun, Cenyuan Zhang, David Lo 等ICSE 2026
- Getting the most out of your tokenizer for pre-training and domain adaptationGautier Dagan, Gabriel Synnaeve, Baptiste RozièreICML 2024 · 被引用 68 次
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li 等ISSTA 2024 · 被引用 8 次
- AST-T5: Structure-Aware Pretraining for Code Generation and UnderstandingLinyuan Gong, Mostafa Elhoushi, Alvin CheungICML 2024 · 被引用 42 次
- Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language ModelsYan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang 等FSE 2024 · 被引用 6 次
