Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code Generation
Yu Yu, Zhihong Sun, Jia Li, Yao Wan, Chuanyi Li, Hongyu Zhang, Ruyun Wang, Tao Huang, Zhi Jin, Ge Li, Chen Lyu
摘要
Large Language Models (LLMs) are capable of generating syntactically correct and functionally complete programs, greatly streamlining software development. However, recent studies reveal that these programs typically execute substantially slower than human-optimized counterparts. Existing approaches to bridging this efficiency gap typically involve either iteratively optimizing code after generation or fine-tuning models on corpora of efficient code. Yet, these methods expose the model to efficiency signals only by mimicking complete, optimized solutions, without explicitly encoding the structural code patterns essential for achieving high runtime performance. Addressing this gap presents two core challenges: (1) extracting and representing latent, efficiency-oriented structural patterns embedded within complex syntax and control flows, and (2) effectively learning these patterns without destabilizing the semantic training of LLMs. To tackle these challenges, we propose EffiSkel, an effi ciency skel eton-guided framework that explicitly extracts and learns efficiency skeletons—abstract, reusable structural patterns underpinning efficient code—by leveraging three complementary strategies: lexical analysis based on token-frequency saliency, syntactic analysis using similarity over Abstract Syntax Trees (ASTs), and dynamic line-level profiling of execution time. These skeletons are integrated into a multi-task learning regime that jointly optimizes code generation and skeleton prediction, introducing an explicit inductive bias toward efficiency-aware code generation. Experiments across multiple programming languages and benchmarks demonstrate that EffiSkel significantly enhances both functional correctness and efficiency, resulting on Mercury with DeepSeek-Coder (6.7B) a +11.11% (vs. EffiCoder) and +3.71% (vs. CodeDPO) higher Efficiency Ratio (ER), and a +0.36 (vs. EffiCoder) and +0.22 (vs. CodeDPO) increase in Average Speedup (AS). These results highlight the effectiveness of explicitly modeling efficiency skeletons in improving the runtime performance of code generated by LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- TreeGen: A Tree-Based Transformer Architecture for Code GenerationZeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun 等AAAI 2020 · 被引用 196 次
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon 等ICLR 2024 · 被引用 141 次
- Exploring and Unleashing the Power of Large Language Models in Automated Code TranslationZhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung 等FSE 2024 · 被引用 72 次
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 被引用 71 次
相关 Paper
- More Than Just Functional: LLM-as-a-Critique for Efficient Code GenerationDerui Zhu, Dingfan Chen, Jinfu Chen, Jens Grossklags 等NeurIPS 2025 · 被引用 2 次
- EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuningDong Huang, Guangtao Zeng, Jianbo Dai, Meng Luo 等ICML 2025
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code GenerationHenrijs Princis, Arindam Sharma, Cristina DavidPLDI 2026
- EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationDong Huang, Jianbo Dai, Han Weng, Puzhen Wu 等NeurIPS 2024 · 被引用 54 次
- How efficient is LLM-generated code? A rigorous & high-standard benchmarkRuizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott 等ICLR 2025
