SemRep : Generative Code Representation Learning with Code Transformations
Weichen Li, Jiamin Song, Bogdan Stoica, Arav Dhoot, Gabriel Ryan, Shengyu Fu, Kexin Pei
Abstract
Code transformation is a foundational capability in the software development process, where its effectiveness relies on constructing a high-quality code representation to characterize the input code semantics and guide the transformation. Existing approaches treat code transformation as an end-to-end learning task, leaving the construction of the representation needed for semantic reasoning implicit in model weights or relying on rigid compiler-level abstractions. We present SemRep, a framework that improves code transformation through generative code representation learning. Our key insight is to employ the semantics-preserving transformations as the intermediate representation, which serves as both a generative mid-training task and the guidance for subsequent instruction-specific code transformations. Across general code editing and optimization tasks (e.g., GPU kernel optimization), SemRep outperforms the extensively finetuned baselines with strictly the same training budget by 6.9% in correctness, 1.1 in performance, 13.9% in generalization, and 6.7% in robustness. With the improved exploration of diverse code transformations, SemRep is particularly amenable to evolutionary search. Combined with an evolutionary coding agent, SemRep finds optimizations that 685B larger-weight baselines fail to discover while achieving the same performance with 25% less inference compute.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionYuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux et al.NeurIPS 2025 · 291 citations
- Chain of Code: Reasoning with a Language Model-Augmented Code EmulatorChengshu Li, Jacky Liang, Andy Zeng, Xinyun Chen et al.ICML 2024 · 155 citations
Related papers
- KernelFoundry: Hardware-Aware Evolutionary GPU Kernel OptimizationNina Wiedemann, Quentin Leboutet, Michael Paulitsch, Diana Wofk et al.ICML 2026 · 11 citations
- Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance OptimizationMohammad Zaeed, Tanzima Z. Islam, Vladimir IndicKDD 2026
- STARK: Strategic Team of Agents for Refining KernelsJuncheng Dong, Yang Yang, Tao Liu, Yang Wang et al.ICLR 2026 · 26 citations
- EditLord: Learning Code Transformation Rules for Code EditingWeichen Li, Albert Jan, Baishakhi Ray, Junfeng Yang et al.ICML 2025
- Exploring Representation-level Augmentation for Code SearchHaochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang et al.EMNLP 2022 · 17 citations
