Towards Learning (Dis)-Similarity of Source Code from Program Contrasts
Yangruibo Ding, Luca Buratti, Saurabh Pujar, Alessandro Morari, Baishakhi Ray, Saikat Chakraborty
摘要
Understanding the functional (dis)-similarity of source code is significant for code modeling tasks such as software vulnerability and code clone detection. We present DISCO (DISsimilarity of COde), a novel self-supervised model focusing on identifying (dis)similar functionalities of source code. Different from existing works, our approach does not require a huge amount of randomly collected datasets. Rather, we design structure-guided code transformation algorithms to generate synthetic code clones and inject real-world security bugs, augmenting the collected datasets in a targeted way. We propose to pre-train the Transformer model with such automatically generated program contrasts to better identify similar code in the wild and differentiate vulnerable programs from benign ones. To better capture the structural features of source code, we propose a new cloze objective to encode the local tree-based context (e.g., parents or sibling nodes). We pre-train our model with a much smaller dataset, the size of which is only 5% of the state-of-the-art models' training datasets, to illustrate the effectiveness of our data augmentation and the pre-training approach. The evaluation shows that, even with much less data, DISCO can still outperform the state-of-the-art models in vulnerability and code clone detection tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- An Empirical Study of Deep Learning Models for Vulnerability DetectionBenjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, Wei LeICSE 2023 · 被引用 107 次
- Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChainMarcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar 等ICLR 2024 · 被引用 39 次
- Towards Efficient Fine-Tuning of Pre-trained Code Models: An Experimental Study and BeyondEnsheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du 等ISSTA 2023 · 被引用 37 次
- TRACED: Execution-aware Pre-training for Source CodeYangruibo Ding, Benjamin Steenhoek, Kexin Pei, Gail E. Kaiser 等ICSE 2024 · 被引用 29 次
- Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsSicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo 等ICSE 2024 · 被引用 27 次
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
相关 Paper
- CONCORD: Clone-Aware Contrastive Learning for Source CodeYangruibo Ding, Saikat Chakraborty, Luca Buratti, Saurabh Pujar 等ISSTA 2023 · 被引用 6 次
- Contrastive Code Representation LearningParas Jain, Ajay Jain, Tianjun Zhang, Pieter Abbeel 等EMNLP 2021 · 被引用 11 次
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang 等AAAI 2024 · 被引用 9 次
- Pre-training by Predicting Program Dependencies for Vulnerability Analysis TasksZhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia 等ICSE 2024 · 被引用 15 次
- Uncovering the Limits of Machine Learning for Automatic Vulnerability DetectionNiklas Risse, Marcel BöhmeUSENIX Security 2024 · 被引用 63 次
