Self-Supervised Query Reformulation for Code Search
Yuetian Mao, Chengcheng Wan, Yuze Jiang, Xiaodong Gu
摘要
Automatic query reformulation is a widely utilized technology for enriching user requirements and enhancing the outcomes of code search. It can be conceptualized as a machine translation task, wherein the objective is to rephrase a given query into a more comprehensive alternative. While showing promising results, training such a model typically requires a large parallel corpus of query pairs (i.e., the original query and a reformulated query) that are confidential and unpublished by online code search engines. This restricts its practicality in software development processes. In this paper, we propose SSQR, a self-supervised query reformulation method that does not rely on any parallel query corpus. Inspired by pre-trained models, SSQR treats query reformulation as a masked language modeling task conducted on an extensive unannotated corpus of queries. SSQR extends T5 (a sequence-to-sequence model based on Transformer) with a new pre-training objective named corrupted query completion (CQC), which randomly masks words within a complete query and trains T5 to predict the masked content. Subsequently, for a given query to be reformulated, SSQR identifies potential locations for expansion and leverages the pretrained T5 model to generate appropriate content to fill these gaps. The selection of expansions is then based on the information gain associated with each candidate. Evaluation results demonstrate that our method outperforms unsupervised baselines significantly and achieves competitive performance compared to supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
- Towards a Better Understanding of Query Reformulation Behavior in Web SearchJia Chen, Jiaxin Mao, Yiqun Liu, Fan Zhang 等WWW 2021 · 被引用 67 次
- Automated Query Reformulation for Efficient Search based on Query Logs From Stack OverflowKaibo Cao, Chunyang Chen, Sebastian Baltes, Christoph Treude 等ICSE 2021 · 被引用 63 次
相关 Paper
- Improving Sequence-to-Sequence Pre-training via Sequence Span RewritingWangchunshu Zhou, Tao Ge, Canwen Xu, Ke Xu 等EMNLP 2021 · 被引用 9 次
- ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement LearningChangtai Zhu, Siyin Wang, Ruijun Feng, Kai Song 等EMNLP 2025 · 被引用 2 次
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio 等ICSE 2021 · 被引用 9 次
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang 等ACL 2022
- ConvGQR: Generative Query Reformulation for Conversational SearchFengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu 等ACL 2023 · 被引用 29 次
