XSearch: Explainable Code Search via Concept-to-Code Alignment
Yiming Liu, Ruofan Liu, Yun Lin, Zicong Zhang, Weiyu Kong, Pengnian Qi, Xiao Cheng, Weinan Zhang, Qianxiang Wang, Linpeng Huang
Abstract
With the emergence of deep learning, semantic code search has been widely adopted in both academia and industry. These approaches embed natural-language queries and code snippets into a shared embedding space and retrieve results based on vector similarity. Despite their strong performance on benchmark datasets, they often suffer from poor explainability and generalization. Retrieved code may appear semantically similar yet miss critical functional requirements of the query, while providing no explanation of why the result was retrieved. Moreover, such failures become more severe under distribution shift, where models struggle to generalize to unseen benchmarks. In this work, we propose XSearch, an intrinsically explainable code search framework. Our key insight is that, by relying on global embedding similarity, all existing retrievers inherently take an inductive view. They learn statistical patterns, rather than truly understand the query's functional requirements. Therefore, we address the problem by reformulating code search as a deductive concept alignment problem. At a high level, XSearch (i) identifies functional concepts in the query and (ii) explicitly aligns them with corresponding code statements. This explain-then-predict design not only produces inherent concept-level explanations, but also mitigates shortcut learning that harms out-of-distribution generalization. We train an encoder with explicit concept-alignment objectives and perform retrieval through explicit matching between query concepts and code statements. Experiments show that, when trained on CodeSearchNet with a small model size (GraphCodeBERT with 125M parameters), XSearch improves performance on out-of-distribution benchmarks from 0.02 to 0.33 (15×) over eight state-of-the-art retrievers, and consistently outperforms both encoder-based and decoder-based baselines with up to 7B parameters. A controlled user study further demonstrates that concept-alignment explanations enable users to accept or reject retrieved results both faster and more accurately. The source code is publicly available at https://github.com/code-philia/Xsearch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa2c9ec7-5ebb-480b-8169-fb71c636661aCited by top-tier papers1
Ask how each one uses itBuilds on40
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Unsupervised Translation of Programming LanguagesBaptiste Rozière, Marie-Anne Lachaux, Lowik Chanussot, Guillaume LampleNeurIPS 2020 · 606 citations
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui et al.EMNLP 2023 · 339 citations
Related papers
- Hedgecode: A Multi-Task Hedging Contrastive Learning Framework for Code SearchGong Chen, Xiaoyuan Xie, Daniel Tang, Qi Xin et al.ICSE 2025
- NS3: Neuro-symbolic Semantic Code SearchShushan Arakelyan, Anna Hakhverdyan, Miltiadis Allamanis, Luis Garcia et al.NeurIPS 2022 · 15 citations
- CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code SearchXiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu et al.EMNLP 2022 · 25 citations
- Generating Explanations to Understand and Repair Embedding-Based Entity AlignmentXiaobin Tian, Zequn Sun, Wei HuICDE 2024 · 6 citations
- UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code RetrievalYang Yang, Li Kuang, Jiakun Liu, Zhongxin Liu et al.ICSE 2026
