Self-Supervised Query Reformulation for Code Search
Yuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong Gu
Abstract
Automatic query reformulation is a widely utilized technology for enriching user requirements and enhancing the outcomes of code search. It can be conceptualized as a machine translation task, wherein the objective is to rephrase a given query into a more comprehensive alternative. While showing promising results, training such a model typically requires a large parallel corpus of query pairs (i.e., the original query and a reformulated query) that are confidential and unpublished by online code search engines. This restricts its practicality in software development processes. In this paper, we propose SSQR, a self-supervised query reformulation method that does not rely on any parallel query corpus. Inspired by pre-trained models, SSQR treats query reformulation as a masked language modeling task conducted on an extensive unannotated corpus of queries. SSQR extends T5 (a sequence-to-sequence model based on Transformer) with a new pre-training objective named corrupted query completion (CQC), which randomly masks words within a complete query and trains T5 to predict the masked content. Subsequently, for a given query to be reformulated, SSQR identifies potential locations for expansion and leverages the pretrained T5 model to generate appropriate content to fill these gaps. The selection of expansions is then based on the information gain associated with each candidate. Evaluation results demonstrate that our method outperforms unsupervised baselines significantly and achieves competitive performance compared to supervised methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge et al.ICSE 2022 · 99 citations
- Towards a Better Understanding of Query Reformulation Behavior in Web SearchJia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang et al.WWW 2021 · 67 citations
- Automated Query Reformulation for Efficient Search based on Query Logs From Stack OverflowKaibo Cao, Chunyang Chen, Sebastian Baltes, Christoph Treude et al.ICSE 2021 · 63 citations
Related papers
- Improving Sequence-to-Sequence Pre-training via Sequence Span RewritingWangchunshu Zhou, Tao Ge, Canwen Xu, Ke Xu et al.EMNLP 2021 · 9 citations
- ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement LearningChangtai Zhu, Siyin Wang, Ruijun Feng, Kai Song et al.EMNLP 2025 · 2 citations
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio et al.ICSE 2021 · 9 citations
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang et al.ACL 2022
- ConvGQR: Generative Query Reformulation for Conversational SearchFengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu et al.ACL 2023 · 29 citations
