Chiseling Out Efficiency: Structured Skeleton Supervision for Efficient Code Generation
Yu Yu, Zhihong Sun, Jia Li, Yao Wan, Chuanyi Li, Hongyu Zhang, Ruyun Wang, Tao Huang, Zhi Jin, Ge Li, Chen Lyu
Abstract
Large Language Models (LLMs) are capable of generating syntactically correct and functionally complete programs, greatly streamlining software development. However, recent studies reveal that these programs typically execute substantially slower than human-optimized counterparts. Existing approaches to bridging this efficiency gap typically involve either iteratively optimizing code after generation or fine-tuning models on corpora of efficient code. Yet, these methods expose the model to efficiency signals only by mimicking complete, optimized solutions, without explicitly encoding the structural code patterns essential for achieving high runtime performance. Addressing this gap presents two core challenges: (1) extracting and representing latent, efficiency-oriented structural patterns embedded within complex syntax and control flows, and (2) effectively learning these patterns without destabilizing the semantic training of LLMs. To tackle these challenges, we propose EffiSkel, an effi ciency skel eton-guided framework that explicitly extracts and learns efficiency skeletons—abstract, reusable structural patterns underpinning efficient code—by leveraging three complementary strategies: lexical analysis based on token-frequency saliency, syntactic analysis using similarity over Abstract Syntax Trees (ASTs), and dynamic line-level profiling of execution time. These skeletons are integrated into a multi-task learning regime that jointly optimizes code generation and skeleton prediction, introducing an explicit inductive bias toward efficiency-aware code generation. Experiments across multiple programming languages and benchmarks demonstrate that EffiSkel significantly enhances both functional correctness and efficiency, resulting on Mercury with DeepSeek-Coder (6.7B) a +11.11% (vs. EffiCoder) and +3.71% (vs. CodeDPO) higher Efficiency Ratio (ER), and a +0.36 (vs. EffiCoder) and +0.22 (vs. CodeDPO) increase in Average Speedup (AS). These results highlight the effectiveness of explicitly modeling efficiency skeletons in improving the runtime performance of code generated by LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1517aa96-d7cf-493c-85ec-7209d491aa90Builds on19
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- TreeGen: A Tree-Based Transformer Architecture for Code GenerationZeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun et al.AAAI 2020 · 196 citations
- Learning Performance-Improving Code EditsAlexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon et al.ICLR 2024 · 141 citations
- Exploring and Unleashing the Power of Large Language Models in Automated Code TranslationZhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung et al.FSE 2024 · 72 citations
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 71 citations
Related papers
- More Than Just Functional: LLM-as-a-Critique for Efficient Code GenerationDerui Zhu, Dingfan Chen, Jinfu Chen, Jens Grossklags et al.NeurIPS 2025 · 2 citations
- EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuningDong Huang, Guangtao Zeng, Jianbo Dai, Meng Luo et al.ICML 2025
- TreeCoder: Systematic Exploration and Optimisation of Decoding and Constraints for LLM Code GenerationHenrijs Princis, Arindam Sharma, Cristina DavidPLDI 2026
- EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationDong Huang, Jianbo Dai, Han Weng, Puzhen Wu et al.NeurIPS 2024 · 54 citations
- How efficient is LLM-generated code? A rigorous & high-standard benchmarkRuizhong Qiu, Weiliang Will Zeng, James Ezick, Christopher Lott et al.ICLR 2025
