Response Enhanced Semi-supervised Dialogue Query Generation
Jianheng Huang, Ante Wang, Linfeng Gao, Linfeng Song, Jinsong Su
Abstract
Leveraging vast and continually updated knowledge from the Internet has been considered an important ability for a dialogue system. Therefore, the dialogue query generation task is proposed for generating search queries from dialogue histories, which will be submitted to a search engine for retrieving relevant websites on the Internet. In this regard, previous efforts were devoted to collecting conversations with annotated queries and training a query producer (QP) via standard supervised learning. However, these studies still face the challenges of data scarcity and domain adaptation. To address these issues, in this paper, we propose a semi-supervised learning framework -- SemiDQG, to improve model performance with unlabeled conversations. Based on the observation that the search query is typically related to the topic of dialogue response, we train a response-augmented query producer (RA) to provide rich and effective training signals for QP. We first apply a similarity-based query selection strategy to select high-quality RA-generated pseudo queries, which are used to construct pseudo instances for training QP and RA. Then, we adopt the REINFORCE algorithm to further enhance QP, with RA-provided rewards as fine-grained training signals. Experimental results and in-depth analysis of three benchmarks show the effectiveness of our framework in cross-domain and low-resource scenarios. Particularly, SemiDQG significantly surpasses ChatGPT and competitive baselines. Our code is available at https://github.com/DeepLearnXMU/SemiDQG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b9b7857-9302-4f32-8801-e6dac2e33bdfCited by top-tier papers1
Ask how each one uses itBuilds on6
- BRIO: Bringing Order to Abstractive SummarizationYixin Liu, Pengfei Liu, Dragomir R. Radev, Graham NeubigACL 2022 · 329 citations
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven ConversationHao Zhou, Chujie Zheng, Kaili Huang, Minlie Huang et al.ACL 2020 · 106 citations
- Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage RetrievalDevang Kulshreshtha, Robert Belfer, Iulian Vlad Serban, Siva ReddyEMNLP 2021 · 11 citations
- Exploring All-In-One Knowledge Distillation Framework for Neural Machine TranslationZhongjian Miao, Wen Zhang, Jinsong Su, Xiang Li et al.EMNLP 2023 · 5 citations
Related papers
- Learning to Relate to Previous Turns in Conversational SearchFengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao et al.KDD 2023 · 16 citations
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 4 citations
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu et al.AAAI 2022 · 181 citations
- Q-TOD: A Query-driven Task-oriented Dialogue SystemXin Tian, Yingzhan Lin, Mengfei Song, Siqi Bao et al.EMNLP 2022 · 13 citations
- KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationQi Zhu, Fei Mi, Zheng Zhang, Yasheng Wang et al.AAAI 2023 · 5 citations
