Natural Is the Best: Model-Agnostic Code Simplification for Pre-trained Large Language Models
Yan Wang, Xiaoning Li, Tien N. Nguyen, Shaohua Wang, Chao Ni, Ling Ding
摘要
Pre-trained Large Language Models (LLM) have achieved remarkable successes in several domains. However, code-oriented LLMs are often heavy in computational complexity, and quadratically with the length of the input code sequence. Toward simplifying the input program of an LLM, the state-of-the-art approach has the strategies to filter the input code tokens based on the attention scores given by the LLM. The decision to simplify the input program should not rely on the attention patterns of an LLM, as these patterns are influenced by both the model architecture and the pre-training dataset. Since the model and dataset are part of the solution domain, not the problem domain where the input program belongs, the outcome may differ when the model is pre-trained on a different dataset. We propose S lim C ode , a model-agnostic code simplification solution for LLMs that depends on the nature of input code tokens. As an empirical study on the LLMs including CodeBERT, CodeT5, and GPT-4 for two main tasks: code search and summarization, we reported that 1) the removal ratio of code has a linear-like relation with the saving ratio on training time, 2) the impact of categorized tokens on code simplification can vary significantly, 3) the impact of categorized tokens on code simplification is task-specific but model-agnostic, and 4) the above findings hold for the paradigm–prompt engineering and interactive in-context learning. The empirical results showed that S lim C ode can improve the state-of-the-art technique by <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline"> mml:mrow mml:mn9.46</mml:mn> mml:mtext%</mml:mtext> </mml:mrow> </mml:math> and <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline"> mml:mrow mml:mn5.15</mml:mn> mml:mtext%</mml:mtext> </mml:mrow> </mml:math> in terms of MRR and BLEU score on code search and summarization, respectively. More importantly, S lim C ode is 133 times faster than the state-of-the-art approach. Additionally, S lim C ode can reduce the cost of invoking GPT-4 by up to <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" display="inline"> mml:mrow mml:mn24</mml:mn> mml:mtext%</mml:mtext> </mml:mrow> </mml:math> per API query, while still producing comparable results to those with the original code. With this result, we call for a new direction on code-based, model-agnostic code simplification solutions to further empower LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu 等ICSE 2025 · 被引用 6 次
- Token Sugar: Making Source Code Sweeter for LLMs through Token-Efficient ShorthandZhensu Sun, Chengran Yang, Xiaoning Du, Zhou Yang 等ASE 2025 · 被引用 2 次
- Seeing Is Coding: On the Effectiveness of Vision Language Models in Code UnderstandingYuling Shi, Chaoxiang Xie, Zhensu Sun, Yeheng Chen 等ISSTA 2026 · 被引用 1 次
- LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language ModelsYan Wang, Ling Ding, Tien N. Nguyen, Shaohua Wang 等ACL 2025 · 被引用 1 次
- Reducing Cost of LLM Agents with Trajectory ReductionYuan-An Xiao, Pengfei Gao, Chao Peng, Yingfei XiongFSE 2026 · 被引用 1 次
它引用的顶会 Paper22
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Learning and Evaluating Contextual Embedding of Source CodeAditya Kanade, Petros Maniatis, Gogul Balakrishnan, Kensen ShiICML 2020 · 被引用 438 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
相关 Paper
- Diet code is healthy: simplifying programs for pre-trained models of codeZhaowei Zhang, Hongyu Zhang, Beijun Shen, Xiaodong GuFSE 2022 · 被引用 39 次
- AI Coders Are among Us: Rethinking Programming Language Grammar towards Efficient Code GenerationZhensu Sun, Xiaoning Du, Zhou Yang, Li Li 等ISSTA 2024 · 被引用 8 次
- What Do They Capture? - A Structural Analysis of Pre-Trained Language Models for Source CodeYao Wan, Wei Zhao, Hongyu Zhang, Yulei Sui 等ICSE 2022 · 被引用 66 次
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- Source Code Summarization in the Era of Large Language ModelsWeisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang 等ICSE 2025 · 被引用 37 次
