Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
Reina Akama, Sho Yokoi, Jun Suzuki, Kentaro Inui
Abstract
Large-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Heterogeneous-Branch Collaborative Learning for Dialogue GenerationYiwei Li, Shaoxiong Feng, Bin Sun, Kan LiAAAI 2023 · 4 citations
- Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service DialoguesMengze Hong, Wailing Ng, Chen Jason Zhang, Yuanfeng Song et al.EMNLP 2025 · 1 citation
- A Textual Dataset for Situated Proactive Response SelectionNaoki Otani, Jun Araki, HyeongSik Kim, Eduard H. HovyACL 2023 · 1 citation
- A Model-agnostic Data Manipulation Method for Persona-based Dialogue GenerationYu Cao, Wei Bi, Meng Fang, Shuming Shi et al.ACL 2022
Builds on3
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- Grounding Conversations with Improvised DialoguesHyundong Cho, Jonathan MayACL 2020 · 27 citations
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 10 citations
Related papers
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang et al.ACL 2021
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou et al.ACL 2020 · 69 citations
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
- Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesZekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng et al.ACL 2021
