Bridge and Hint: Extending Pre-trained Language Models for Long-Range Code
Yujia Chen, Cuiyun Gao, Zezhou Yang, Hongyu Zhang, Qing Liao
Abstract
In the field of code intelligence, effectively modeling long-range code poses a significant challenge. Existing pre-trained language models (PLMs) such as UniXcoder have achieved remarkable success, but they still face difficulties with long code inputs. This is mainly due to their limited capacity to maintain contextual continuity and memorize the key information over long-range code. To alleviate the difficulties, we propose EXPO, a framework for EXtending Pre-trained language models for lOng-range code. EXPO incorporates two innovative memory mechanisms we propose in this paper: Bridge Memory and Hint Memory. Bridge Memory uses a tagging mechanism to connect disparate snippets of long-range code, helping the model maintain contextual coherence. Hint Memory focuses on crucial code elements throughout the global context, such as package imports, by integrating a 𝑘NN attention layer to adaptively select the relevant code elements. This dual-memory approach bridges the gap between understanding local code snippets and maintaining global code coherence, thereby enhancing the model’s overall comprehension of long code sequences. We validate the effectiveness of EXPO on five popular pre-trained language models such as UniXcoder and two code intelligence tasks including API recommendation and vulnerability detection. Experimental results demonstrate that EXPO significantly improves the pre-training language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00003849-d679-485d-ad01-42316153bcd4Cited by top-tier papers2
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia et al.ICSE 2026 · 1 citation
- Safe4U: Identifying Unsound Safe Encapsulations of Unsafe Calls in Rust using LLMsHuan Li, Bei Wang, Xing Hu, Xin XiaISSTA 2025 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 283 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
Related papers
- CTX-Coder: Cross-Attention Architectures Empower LLMs for Long-Context Vulnerability DetectionJujie Wang, Kangfeng Zheng, Bin Wu, Chunhua Wu et al.AAAI 2026
- LongCoder: A Long-Range Pre-trained Language Model for Code CompletionDaya Guo, Canwen Xu, Nan Duan, Jian Yin et al.ICML 2023 · 150 citations
- UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationDaya Guo, Shuai Lu, Nan Duan, Yanlin Wang et al.ACL 2022
- LOSVER: Line-Level Modifiability Signal-Guided Vulnerability Detection and ClassificationDoha Nam, Jongmoon BaikASE 2025
- SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability DetectionXin-Cheng Wen, Cuiyun Gao, Shuzheng Gao, Yang Xiao et al.ISSTA 2024 · 17 citations
