UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code Retrieval
Yang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu, Yingjie Xia, David Lo
摘要
Effective code retrieval is indispensable and it has become an important paradigm to search code in hybrid mode using both natural language and code snippets. Nevertheless, it remains unclear whether existing approaches can effectively leverage such hybrid queries, particularly in cross-language contexts. We conduct a comprehensive empirical study of representative code models and reveal three challenges: (1) insufficient semantic understanding; (2) inefficient fusion in hybrid code retrieval; and (3) weak generalization in cross-language scenarios. To address these challenges, we propose UniCoR, a novel self-supervised framework designed to learn Unified Code Representations that are semantically robust, modally collaborative, and language-agnostic. Firstly, we design a multi-perspective supervised contrastive learning module to enhance semantic understanding and modality fusion. It aligns representations from multiple perspectives, including code-to-code, natural language-to-code, and natural language-to-natural language, enforcing the model to capture a semantic essence among modalities. Secondly, we introduce a representation distribution consistency learning module to improve cross-language generalization, which explicitly aligns the feature distributions of different programming languages, enabling language-agnostic representation learning. Extensive experiments on both an empirical benchmark and a large-scale benchmark show that UniCoR outperforms all baseline models, achieving an average improvement of 8.64% in MRR and 11.54% in MAP over the best-performing baseline. Furthermore, UniCoR exhibits stability in hybrid code retrieval and generalization capability in cross-language scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- How could Neural Networks understand Programs?Dinglan Peng, Shuxin Zheng, Yatao Li, Guolin Ke 等ICML 2021 · 被引用 75 次
- CoCoSoDa: Effective Contrastive Learning for Code SearchEnsheng Shi, Yanlin Wang, Wenchao Gu, Lun Du 等ICSE 2023 · 被引用 45 次
- Cross-Domain Deep Code Search with Meta LearningYitian Chai, Hongyu Zhang, Beijun Shen, Xiaodong GuICSE 2022 · 被引用 42 次
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang 等ACL 2024 · 被引用 21 次
相关 Paper
- Self-Supervised Contrastive Learning for Code Retrieval and Summarization via Semantic-Preserving TransformationsNghi D. Q. Bui, Yijun Yu, Lingxiao JiangSIGIR 2021 · 被引用 98 次
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu 等EMNLP 2022 · 被引用 25 次
- UniXcoder: Unified Cross-Modal Pre-training for Code RepresentationDaya Guo, Shuai Lu, Nan Duan, Yanlin Wang 等ACL 2022
- Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchGong Chen, Xiaoyuan Xie, Daniel Tang, Qi Xin 等ICSE 2025
- UNICS: Multilingual Code Search via Unified Pseudocode and Contrastive Transfer LearningYe Fan, Jidong Ge, Chuanyi Li, LiGuo Huang 等FSE 2026
