Beyond Sequences: Two-dimensional Representation and Dependency Encoding for Code Generation
Xiangyu Zhang, Yu Zhou, Guang Yang, Wei Cheng, Taolue Chen
Abstract
The advent of large language models has significantly advanced automatic code generation, transforming the way programmers writing code. Inspired by natural language processing, mainstream code generation approaches represent code as a linear sequence of tokens. In this paper, we propose to represent code snippets as two-dimensional entities, where both code lines and tokens within lines are explicitly modeled. This representation allows us to capture the hierarchical and spatial structure of code, especially the dependencies between code lines. Our method CoDE introduces a dependency encoding approach that leverages dictionary learning to perform semantic matching between code lines. As such, it avoids the reliance on strict position indices, leading to better generalization to code with diverse context and lengths. We thoroughly evaluate CoDE based on four categories of tasks. The experimental results showcase its generalizability, context understanding and retrieval, as well as interpretability in code generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d47d051d-8652-49a6-bf8c-244e1fbc5e18Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
Related papers
- Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary LearningJohn Wu, David Wu, Jimeng SunEMNLP 2024 · 4 citations
- Alignment with Fill-In-the-Middle for Enhancing Code GenerationHouxing Ren, Zimu Lu, Weikang Shi, Haotian Hou et al.EMNLP 2025
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang et al.ICSE 2024 · 124 citations
- Hot or Cold? Adaptive Temperature Sampling for Code Generation with Large Language ModelsYuqi Zhu, Jia Li, Ge Li, Yunfei Zhao et al.AAAI 2024 · 68 citations
- One Size Does Not Fit All: Revisiting Code Context Engineering for Repository-Level Code GenerationYichen Li, Qiye Lin, Yun Peng, Zhihan Jiang et al.FSE 2026
