Filtering Noisy Dialogue Corpora by Connectivity and Content Relatedness
Reina Akama, Sho Yokoi, Jun Suzuki, Kentaro Inui
摘要
Large-scale dialogue datasets have recently become available for training neural dialogue agents. However, these datasets have been reported to contain a non-negligible number of unacceptable utterance pairs. In this paper, we propose a method for scoring the quality of utterance pairs in terms of their connectivity and relatedness. The proposed scoring method is designed based on findings widely shared in the dialogue and linguistics research communities. We demonstrate that it has a relatively good correlation with the human judgment of dialogue quality. Furthermore, the method is applied to filter out potentially unacceptable utterance pairs from a large-scale noisy dialogue corpus to ensure its quality. We experimentally confirm that training data filtered by the proposed method improves the quality of neural dialogue agents in response generation. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Heterogeneous-Branch Collaborative Learning for Dialogue GenerationYiwei Li, Shaoxiong Feng, Bin Sun, Kan LiAAAI 2023 · 被引用 4 次
- Dial-In LLM: Human-Aligned LLM-in-the-loop Intent Clustering for Customer Service DialoguesMengze Hong, Wailing Ng, Chen Jason Zhang, Yuanfeng Song 等EMNLP 2025 · 被引用 1 次
- A Textual Dataset for Situated Proactive Response SelectionNaoki Otani, Jun Araki, HyeongSik Kim, Eduard H. HovyACL 2023 · 被引用 1 次
- A Model-agnostic Data Manipulation Method for Persona-based Dialogue GenerationYu Cao, Wei Bi, Meng Fang, Shuming Shi 等ACL 2022
它引用的顶会 Paper3
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- Grounding Conversations with Improvised DialoguesHyundong Cho, Jonathan MayACL 2020 · 被引用 27 次
- USR: An Unsupervised and Reference Free Evaluation Metric for Dialog GenerationShikib Mehri, Maxine EskénaziACL 2020 · 被引用 10 次
相关 Paper
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
- DynaEval: Unifying Turn and Dialogue Level EvaluationChen Zhang, Yiming Chen, Luis Fernando D'Haro, Yan Zhang 等ACL 2021
- Towards Holistic and Automatic Evaluation of Open-Domain Dialogue GenerationBo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou 等ACL 2020 · 被引用 69 次
- Proxy Indicators for the Quality of Open-domain DialoguesRostislav Nedelchev, Jens Lehmann, Ricardo UsbeckEMNLP 2021
- Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue UtterancesZekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng 等ACL 2021
