Exploring Representation-level Augmentation for Code Search
Haochen Li, Chunyan Miao, Cyril Leung, Yanxian Huang, Yuan Huang, Hongyu Zhang, Yanlin Wang
摘要
Code search, which aims at retrieving the most relevant code fragment for a given natural language query, is a common activity in software development practice. Recently, contrastive learning is widely used in code search research, where many data augmentation approaches for source code (e.g., semantic-preserving program transformation) are proposed to learn better representations. However, these augmentations are at the raw-data level, which requires additional code analysis in the preprocessing stage and additional training costs in the training stage. In this paper, we explore augmentation methods that augment data (both code and query) at representation level which does not require additional data processing and training, and based on this we propose a general format of representationlevel augmentation that unifies existing methods. Then, we propose three new augmentation methods (linear extrapolation, binary interpolation, and Gaussian scaling) based on the general format. Furthermore, we theoretically analyze the advantages of the proposed augmentation methods over traditional contrastive learning methods on code search. We experimentally evaluate the proposed representationlevel augmentation methods with state-of-theart code search models on a large-scale public dataset consisting of six programming languages. The experimental results show that our approach can consistently boost the performance of the studied code search models. Our source code is available at https://github. com/Alex-HaochenLi/RACS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- iSMELL: Assembling LLMs with Expert Toolsets for Code Smell Detection and RefactoringDi Wu, Fangwen Mu, Lin Shi, Zhaoqiang Guo 等ASE 2024 · 被引用 19 次
- CONSIDER: Commonalities and Specialties Driven Multilingual Code Retrieval FrameworkRui Li, Liyang He, Qi Liu, Yuze Zhao 等AAAI 2024 · 被引用 11 次
- XSearch: Explainable Code Search via Concept-to-Code AlignmentYiming Liu, Ruofan Liu, Yun Lin, Zicong Zhang 等ISSTA 2026 · 被引用 1 次
- MGS3: A Multi-Granularity Self-Supervised Code Search FrameworkRui Li, Junfeng Kang, Qi Liu, Liyang He 等KDD 2025
- GiFT: Gibbs Fine-Tuning for Code GenerationHaochen Li, Wanjin Feng, Xin Zhou, Zhiqi ShenACL 2025
它引用的顶会 Paper10
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Towards Domain-Agnostic Contrastive LearningVikas Verma, Thang Luong, Kenji Kawaguchi, Hieu Pham 等ICML 2021 · 被引用 131 次
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 被引用 98 次
相关 Paper
- Uncertainty-Aware Contrastive Learning with Hard Negative Sampling for Code Search TasksHan Liu, Jiaqing Zhan, Qin ZhangAAAI 2025 · 被引用 1 次
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du 等ICSE 2023 · 被引用 45 次
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu 等EMNLP 2022 · 被引用 25 次
- Code Representation Learning at ScaleDejiao Zhang, Wasi Uddin Ahmad, Ming Tan, Hantian Ding 等ICLR 2024 · 被引用 32 次
- ContraBERT: Enhancing Code Pre-trained Models via Contrastive LearningShangqing Liu, Bozhi Wu, Xiaofei Xie, Guozhu Meng 等ICSE 2023 · 被引用 56 次
