Contextuality of Code Representation Learning
Yi Li, Shaohua Wang, Tien N. Nguyen
Abstract
Advanced machine learning models (ML) have been successfully leveraged in several software engineering (SE) applications. The existing SE techniques have used the embedding models ranging from static to contextualized ones to build the vectors for program units. The contextualized vectors address a phenomenon in natural language texts called polysemy, which is the coexistence of different meanings of a word/phrase. However, due to different nature, program units exhibit the nature of mixed polysemy. Some code tokens and statements exhibit polysemy while other tokens (e.g., keywords, separators, and operators) and statements maintain the same meaning in different contexts. A natural question is whether static or contextualized embeddings fit better with the nature of mixed polysemy in source code. The answer to this question is helpful for the SE researchers in selecting the right embedding model. We conducted experiments on 12 popular sequence-/tree-/graph-based embedding models and on the samples of a dataset of 10,222 Java projects with +14M methods. We present several contextuality evaluation metrics adapted from natural-language texts to code structures to evaluate the embeddings from those models. Among several findings, we found that the models with higher contextuality help a bug detection model perform better than the static ones. Neither static nor contextualized embedding models fit well with the mixed polysemy nature of source code. Thus, we develop Hycode, a hybrid embedding model that fits better with the nature of mixed polysemy in source code.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7eee8460-64d2-46a9-b9bf-0f17f5273fd6Cited by top-tier papers2
- Your Fix Is My Exploit: Enabling Comprehensive DL Library API Fuzzing with Large Language ModelsKunpeng Zhang, Shuai Wang, Jitao Han, Xiaogang Zhu et al.ICSE 2025 · 6 citations
- LLM-Powered Silent Bug Fuzzing in Deep Learning Libraries via Versatile and Controlled Bug TransferKunpeng Zhang, Dongwei Xiao, Daoyuan Wu, Shuai Wang et al.OOPSLA 2026 · 1 citation
Related papers
- Semantic bug seeding: a learning-based approach for creating realistic bugsJibesh Patra, Michael PradelFSE 2021 · 63 citations
- UTANGO: untangling commits with context-aware, graph-based, code change clustering learning modelYi Li, Shaohua Wang, Tien N. NguyenFSE 2022 · 16 citations
- Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?Yu Zhao, Lina Gong, Zhiqiu Huang, Yongwei Wang et al.ASE 2024 · 10 citations
- A Knowledge Enhanced Large Language Model for Bug LocalizationYue Li, Bohan Liu, Ting Zhang, Zhiqi Wang et al.FSE 2025 · 3 citations
- Learning semantic program embeddings with graph interval neural networkYu Wang, Ke Wang, Fengjuan Gao, Linzhang WangOOPSLA 2020 · 61 citations
