Rep2Vec: Repository Embedding via Heterogeneous Graph Adversarial Contrastive Learning
Yiyue Qian, Yiming Zhang, Qianlong Wen, Yanfang Ye, Chuxu Zhang
摘要
Driven by the exponential increase of software and the advent of the pull-based development system Git, a large amount of open-source software has emerged on various social coding platforms. GitHub, as the largest platform, not only attracts developers and researchers to contribute legitimate software and research-related source code but has also become a popular platform for an increasing number of cybercriminals to perform continuous cyberattacks. Hence, some tools have been developed to learn representations of repositories on GitHub for various related applications (e.g., malicious repository detection) recently. However, most of them merely focus on code content while ignoring the rich relational data among repositories. In addition, they usually require a mass of resources to obtain sufficient labeled data for model training while ignoring the usefully handy unlabeled data. To this end, we propose a novel model Rep2Vec which integrates the code content, the structural relations, and the unlabeled data to learn the repository representations. First, to comprehensively model the repository data, we build a repository heterogeneous graph (Rep-HG) which is encoded by a graph neural network. Afterwards, to fully exploit unlabeled data in Rep-HG, we introduce adversarial attacks to generate more challenging contrastive pairs for the contrastive learning module to train the encoder in node view and meta-path view simultaneously. To alleviate the workload of the encoder against attacks, we further design a dual-stream contrastive learning module that integrates contrastive learning on adversarial graph and original graph together. Finally, the pre-trained encoder is fine-tuned to the downstream task, and further enhanced by a knowledge distillation module. Extensive experiments on the collected dataset from GitHub demonstrate the effectiveness of Rep2Vec in comparison with state-of-the-art methods for multiple repository tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Co-Modality Graph Contrastive Learning for Imbalanced Node ClassificationYiyue Qian, Chunhui Zhang, Yiming Zhang, Qianlong Wen 等NeurIPS 2022 · 被引用 54 次
- Pre-training Code Representation with Semantic Flow Graph for Effective Bug LocalizationYali Du, Zhongxing YuFSE 2023 · 被引用 18 次
- Adversarial Contrastive Graph Augmentation with Counterfactual RegularizationTao Long, Lei Zhang, Liang Zhang, Laizhong CuiAAAI 2025 · 被引用 5 次
相关 Paper
- Unsupervised Graph Poisoning Attack via Contrastive Loss Back-propagationSixiao Zhang, Hongxu Chen, Xiangguo Sun, Yicong Li 等WWW 2022 · 被引用 52 次
- Dual Space Graph Contrastive LearningHaoran Yang, Hongxu Chen, Shirui Pan, Lin Li 等WWW 2022 · 被引用 58 次
- paper2repo: GitHub Repository Recommendation for Academic PapersHuajie Shao, Dachun Sun, Jiahao Wu, Zecheng Zhang 等WWW 2020 · 被引用 36 次
- Graph Contrastive Backdoor AttacksHangfan Zhang, Jinghui Chen, Lu Lin, Jinyuan Jia 等ICML 2023 · 被引用 25 次
- Adversarial Contrastive Graph Masked AutoEncoder Against Graph Structure and Feature Dual AttacksWeixuan Shen, Xiaobo Shen, Shirui PanAAAI 2025 · 被引用 1 次
