XSearch: Explainable Code Search via Concept-to-Code Alignment
Yiming Liu, Ruofan Liu, Yun Lin, Zicong Zhang, Weiyu Kong, Pengnian Qi, Xiao Cheng, Weinan Zhang, Qianxiang Wang, Linpeng Huang
摘要
With the emergence of deep learning, semantic code search has been widely adopted in both academia and industry. These approaches embed natural-language queries and code snippets into a shared embedding space and retrieve results based on vector similarity. Despite their strong performance on benchmark datasets, they often suffer from poor explainability and generalization. Retrieved code may appear semantically similar yet miss critical functional requirements of the query, while providing no explanation of why the result was retrieved. Moreover, such failures become more severe under distribution shift, where models struggle to generalize to unseen benchmarks. In this work, we propose XSearch, an intrinsically explainable code search framework. Our key insight is that, by relying on global embedding similarity, all existing retrievers inherently take an inductive view. They learn statistical patterns, rather than truly understand the query's functional requirements. Therefore, we address the problem by reformulating code search as a deductive concept alignment problem. At a high level, XSearch (i) identifies functional concepts in the query and (ii) explicitly aligns them with corresponding code statements. This explain-then-predict design not only produces inherent concept-level explanations, but also mitigates shortcut learning that harms out-of-distribution generalization. We train an encoder with explicit concept-alignment objectives and perform retrieval through explicit matching between query concepts and code statements. Experiments show that, when trained on CodeSearchNet with a small model size (GraphCodeBERT with 125M parameters), XSearch improves performance on out-of-distribution benchmarks from 0.02 to 0.33 (15×) over eight state-of-the-art retrievers, and consistently outperforms both encoder-based and decoder-based baselines with up to 7B parameters. A controlled user study further demonstrates that concept-alignment explanations enable users to accept or reject retrieved results both faster and more accurately. The source code is publicly available at https://github.com/code-philia/Xsearch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper40
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 被引用 606 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchGong Chen, Xiaoyuan Xie, Daniel Tang, Qi Xin 等ICSE 2025
- NS3: Neuro-symbolic Semantic Code SearchShushan Arakelyan, Anna Hakhverdyan, Miltiadis Allamanis, Luis Garcia 等NeurIPS 2022 · 被引用 15 次
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu 等EMNLP 2022 · 被引用 25 次
- Generating Explanations to Understand and Repair Embedding-Based Entity AlignmentXiaobin Tian, Zequn Sun, Wei HuICDE 2024 · 被引用 6 次
- UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code RetrievalYang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu 等ICSE 2026
