Lune

ISSTA2026顶会

XSearch: Explainable Code Search via Concept-to-Code Alignment

Yiming Liu, Ruofan Liu, Yun Lin, Zicong Zhang, Weiyu Kong, Pengnian Qi, Xiao Cheng, Weinan Zhang, Qianxiang Wang, Linpeng Huang

2026年份
1被引次数
1顶会引用

摘要

With the emergence of deep learning, semantic code search has been widely adopted in both academia and industry. These approaches embed natural-language queries and code snippets into a shared embedding space and retrieve results based on vector similarity. Despite their strong performance on benchmark datasets, they often suffer from poor explainability and generalization. Retrieved code may appear semantically similar yet miss critical functional requirements of the query, while providing no explanation of why the result was retrieved. Moreover, such failures become more severe under distribution shift, where models struggle to generalize to unseen benchmarks. In this work, we propose XSearch, an intrinsically explainable code search framework. Our key insight is that, by relying on global embedding similarity, all existing retrievers inherently take an inductive view. They learn statistical patterns, rather than truly understand the query's functional requirements. Therefore, we address the problem by reformulating code search as a deductive concept alignment problem. At a high level, XSearch (i) identifies functional concepts in the query and (ii) explicitly aligns them with corresponding code statements. This explain-then-predict design not only produces inherent concept-level explanations, but also mitigates shortcut learning that harms out-of-distribution generalization. We train an encoder with explicit concept-alignment objectives and perform retrieval through explicit matching between query concepts and code statements. Experiments show that, when trained on CodeSearchNet with a small model size (GraphCodeBERT with 125M parameters), XSearch improves performance on out-of-distribution benchmarks from 0.02 to 0.33 (15×) over eight state-of-the-art retrievers, and consistently outperforms both encoder-based and decoder-based baselines with up to 7B parameters. A controlled user study further demonstrates that concept-alignment explanations enable users to accept or reject retrieved results both faster and more accurately. The source code is publicly available at https://github.com/code-philia/Xsearch.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper40

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖