Joint Geometrical and Statistical Domain Adaptation for Cross-domain Code Vulnerability Detection
Qianjin Du, Shiji Zhou, Xiaohui Kuang, Gang Zhao, Jidong Zhai
摘要
In code vulnerability detection tasks, a detector trained on a label-rich source domain fails to provide accurate prediction on new or unseen target domains due to the lack of labeled training data on target domains. Previous studies mainly utilize domain adaptation to perform cross-domain vulnerability detection. But they ignore the negative effect of private semantic characteristics of the target domain for domain alignment, which easily causes the problem of negative transfer. In addition, these methods forcibly reduce the distribution discrepancy between domains and do not take into account the interference of irrelevant target instances for distributional domain alignment, which leads to the problem of excessive alignment. To address the above issues, we propose a novel cross-domain code vulnerability detection framework named MN-CRI. Specifically, we introduce mutual nearest neighbor contrastive learning to align the source domain and target domain geometrically, which could align the common semantic characteristics of two domains and separate out the private semantic characteristics of each domain. Furthermore, we introduce an instance re-weighting scheme to alleviate the problem of excessive alignment. This scheme dynamically assign different weights to instances, reducing the contribution of irrelevant instances so as to achieve better domain alignment. Finally, extensive experiments demonstrate that MNCRI significantly outperforms state-of-the-art cross-domain code vulnerability detection methods by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
相关 Paper
- Cross-Domain Graph Anomaly Detection via Anomaly-Aware Contrastive AlignmentQizhou Wang, Guansong Pang, Mahsa Salehi, Wray L. Buntine 等AAAI 2023 · 被引用 51 次
- Cross-Domain Detection via Graph-Induced Prototype AlignmentMinghao Xu, Hang Wang, Bingbing Ni, Qi Tian 等CVPR 2020
- CL3D: Unsupervised Domain Adaptation for Cross-LiDAR 3D DetectionXidong Peng, Xinge Zhu, Yuexin MaAAAI 2023 · 被引用 37 次
- Mutual Nearest Neighbor Contrast and Hybrid Prototype Self-Training for Universal Domain AdaptationLiang Chen, Qianjin Du, Yihang Lou, Jianzhong He 等AAAI 2022 · 被引用 33 次
- Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-trainingYihan Zeng, Chunwei Wang, Yunbo Wang, Hang Xu 等NeurIPS 2021 · 被引用 36 次
