Robustness Testing of Language Understanding in Task-Oriented Dialog
Jiexi Liu, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan, Hongguang Li, Weiran Nie, Cheng Li, Wei Peng, Minlie Huang
摘要
Most language understanding models in taskoriented dialog systems are trained on a small amount of annotated training data, and evaluated in a small set from the same distribution. However, these models can lead to system failure or undesirable output when being exposed to natural language perturbation or variation in practice. In this paper, we conduct comprehensive evaluation and analysis with respect to the robustness of natural language understanding models, and introduce three important aspects related to language understanding in realworld dialog systems, namely, language variety, speech characteristics, and noise perturbation. We propose a model-agnostic toolkit LAUG to approximate natural language perturbations for testing the robustness issues in taskoriented dialog. Four data augmentation approaches covering the three aspects are assembled in LAUG, which reveals critical robustness issues in state-of-the-art models. The augmented dataset through LAUG can be used to facilitate future research on the robustness testing of language understanding in task-oriented dialog. Recently task-oriented dialog systems have been attracting more and more research efforts (Gao et al., 2019; Zhang et al., 2020b), where understanding user utterances is a critical precursor to the success of such dialog systems. While modern neural networks have achieved state-of-the-art results on language understanding (LU) (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue SystemsHarrison Lee, Raghav Gupta, Abhinav Rastogi, Yuan Cao 等AAAI 2022 · 被引用 40 次
- Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition ErrorsMarek Kubis, Pawel Skórzewski, Marcin Sowanski, Tomasz ZietkiewiczEMNLP 2023 · 被引用 7 次
- Measuring the Effect of Transcription Noise on Downstream Language Understanding TasksOri Shapira, Shlomo E. Chazan, Amir David Nissan CohenACL 2025 · 被引用 3 次
- DAMP: Doubly Aligned Multilingual Parser for Task-Oriented DialogueWilliam Barr Held, Christopher Hidey, Fei Liu, Eric Zhu 等ACL 2023 · 被引用 2 次
- Log-FGAER: Logic-Guided Fine-Grained Address Entity Recognition from Multi-Turn Spoken DialogueXue Han, Yitong Wang, Qian Hu, Pengwei Hu 等EMNLP 2023 · 被引用 1 次
它引用的顶会 Paper5
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 被引用 210 次
- Multi-Task Self-Supervised Learning for Disfluency DetectionShaolei Wang, Wanxiang Che, Qi Liu, Pengda Qin 等AAAI 2020 · 被引用 56 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
相关 Paper
- Self-training Improves Pre-training for Few-shot Learning in Task-oriented Dialog SystemsFei Mi, Wanhao Zhou, Lingjing Kong, Fengyu Cai 等EMNLP 2021 · 被引用 18 次
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong 等ACL 2026
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu 等ACL 2022
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingYingmei Guo, Linjun Shou, Jian Pei, Ming Gong 等EMNLP 2021 · 被引用 2 次
- MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic DialoguesKuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu 等AAAI 2025 · 被引用 10 次
