Cross-Domain Deep Code Search with Meta Learning
Yitian Chai, Hongyu Zhang, Beijun Shen, Xiaodong Gu
摘要
Recently, pre-trained programming language models such as Code-BERT have demonstrated substantial gains in code search. Despite showing great performance, they rely on the availability of large amounts of parallel data to fine-tune the semantic mappings between queries and code. This restricts their practicality in domainspecific languages that have relatively scarce and expensive data. In this paper, we propose CroCS, a novel approach for domainspecific code search. CroCS employs a transfer learning framework where an initial program representation model is pre-trained on a large corpus of common programming languages (such as Java and Python), and is further adapted to domain-specific languages such as Solidity and SQL. Unlike cross-language CodeBERT, which is directly fine-tuned in the target language, CroCS adapts a few-shot meta-learning algorithm called MAML to learn the good initialization of model parameters, which can be best reused in a domain-specific language. We evaluate the proposed approach on two domain-specific languages, namely Solidity and SQL, with model transferred from two widely used languages (Python and Java). Experimental results show that CroCS significantly outperforms conventional pre-trained code models that are directly finetuned in domain-specific languages, and it is particularly effective for scarce data. CCS CONCEPTS • Software and its engineering → Reusability; Automatic programming.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- No more fine-tuning? an experimental evaluation of prompt tuning in code intelligenceChaozheng Wang, Yuanhang Yang, Cuiyun Gao, Yun Peng 等FSE 2022 · 被引用 148 次
- Self-Supervised Query Reformulation for Code SearchYuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong GuFSE 2023 · 被引用 14 次
- Predicting Configuration Performance in Multiple Environments with Sequential Meta-LearningJingzhi Gong, Tao ChenFSE 2024 · 被引用 13 次
- AdaCCD: Adaptive Semantic Contrasts Discovery Based Cross Lingual Adaptation for Code Clone DetectionYangkai Du, Tengfei Ma, Lingfei Wu, Xuhong Zhang 等AAAI 2024 · 被引用 9 次
- Statistical Type Inference for Incomplete ProgramsYaohui Peng, Jing Xie, Qiongling Yang, Hanwen Guo 等FSE 2023 · 被引用 2 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- OCoR: An Overlapping-Aware Code RetrieverQihao Zhu, Zeyu Sun, Xiran Liang, Yingfei Xiong 等ASE 2020 · 被引用 28 次
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio 等ICSE 2021 · 被引用 9 次
相关 Paper
- UNICS: Multilingual Code Search via Unified Pseudocode and Contrastive Transfer LearningYe Fan, Jidong Ge, Chuanyi Li, LiGuo Huang 等FSE 2026
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du 等ICSE 2023 · 被引用 45 次
- Zero-Shot Cross-Domain Code Search without Fine-TuningKeyu Liang, Zhongxin Liu, Chao Liu, Zhiyuan Wan 等FSE 2025 · 被引用 2 次
- Bridging Pre-trained Models and Downstream Tasks for Source Code UnderstandingDeze Wang, Zhouyang Jia, Shanshan Li, Yue Yu 等ICSE 2022 · 被引用 68 次
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 被引用 19 次
