CycleKQR: Unsupervised Bidirectional Keyword-Question Rewriting
Andrea Iovine, Anjie Fang, Besnik Fetahu, Jie Zhao, Oleg Rokhlenko, Shervin Malmasi
Abstract
Users expect their queries to be answered by search systems, regardless of the query's surface form, which include keyword queries and natural questions. Natural Language Understanding (NLU) components of Search and QA systems may fail to correctly interpret semantically equivalent inputs if this deviates from how the system was trained, leading to suboptimal understanding capabilities. We propose the keyword-question rewriting task to improve query understanding capabilities of NLU systems for all surface forms. To achieve this, we present CycleKQR, an unsupervised approach, enabling effective rewriting between keyword and question queries using non-parallel data. Empirically we show the impact on QA performance of unfamiliar query forms for open domain and Knowledge Base QA systems (trained on either keywords or natural language questions). We demonstrate how CycleKQR significantly improves QA performance by rewriting queries into the appropriate form, while at the same time retaining the original semantic meaning of input queries, allowing CycleKQR to improve performance by up to 3% over supervised baselines. Finally, we release a dataset of 66k keyword-question pairs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42f53748-1ff2-4dc1-a2c3-3c021abdbc66Cited by top-tier papers1
Ask how each one uses itBuilds on6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question AnsweringXi Ye, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou et al.ACL 2022 · 203 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- CycleNER: An Unsupervised Training Approach for Named Entity RecognitionAndrea Iovine, Anjie Fang, Besnik Fetahu, Oleg Rokhlenko et al.WWW 2022 · 19 citations
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 9 citations
Related papers
- CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement LearningZeqiu Wu, Yi Luan, Hannah Rashkin, David Reitter et al.EMNLP 2022 · 37 citations
- How to Ask Better Questions? A Large-Scale Multi-Domain Dataset for Rewriting Ill-Formed QuestionsZewei Chu, Mingda Chen, Jing Chen, Miaosen Wang et al.AAAI 2020 · 22 citations
- ICR: Iterative Clarification and Rewriting for Conversational SearchZhiyu Cao, Peifeng Li, Qiaoming ZhuEMNLP 2025
- ConvGQR: Generative Query Reformulation for Conversational SearchFengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu et al.ACL 2023 · 29 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
