GOLD: Improving Out-of-Scope Detection in Dialogues using Data Augmentation
Derek Chen, Zhou Yu
Abstract
Practical dialogue systems require robust methods of detecting out-of-scope (OOS) utterances to avoid conversational breakdowns and related failure modes. Directly training a model with labeled OOS examples yields reasonable performance, but obtaining such data is a resource-intensive process. To tackle this limited-data problem, previous methods focus on better modeling the distribution of in-scope (INS) examples. We introduce GOLD as an orthogonal technique that augments existing data to train better OOS detectors operating in low-data regimes. GOLD generates pseudo-labeled candidates using samples from an auxiliary dataset and keeps only the most beneficial candidates for training through a novel filtering mechanism. In experiments across three target benchmarks, the top GOLD model outperforms all existing methods on all key metrics, achieving relative gains of 52.4%, 48.9% and 50.3% against median baseline performance. We also analyze the unique properties of OOS data to identify key factors for optimally applying our proposed method. 1 1 All code and data for major experiments are available at https://github.com/asappresearch/gold
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Delving into Out-of-Distribution Detection with Vision-Language RepresentationsYifei Ming, Ziyang Cai, Jiuxiang Gu, Yiyou Sun et al.NeurIPS 2022 · 308 citations
- POEM: Out-of-Distribution Detection with Posterior SamplingYifei Ming, Ying Fan, Yixuan LiICML 2022 · 151 citations
- SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationHyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West et al.EMNLP 2023 · 60 citations
- Estimating Soft Labels for Out-of-Domain Intent DetectionHao Lang, Yinhe Zheng, Jian Sun, Fei Huang et al.EMNLP 2022 · 12 citations
- Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain DetectionRheeya Uppaal, Junjie Hu, Yixuan LiACL 2023 · 9 citations
Builds on7
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Self-Supervised Learning for Generalizable Out-of-Distribution DetectionSina Mohseni, Mandar Pitale, J. B. S. Yadawa, Zhangyang WangAAAI 2020 · 229 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional GenerationRuibo Liu, Guangxuan Xu, Chenyan Jia, Weicheng Ma et al.EMNLP 2020 · 62 citations
Related papers
- Out-of-Scope Intent Detection with Self-Supervision and Discriminative TrainingLi-Ming Zhan, Haowen Liang, Bo Liu, Lu Fan et al.ACL 2021
- Auxiliary Prompt Tuning of Vision-Language Models for Few-Shot Out-of-Distribution DetectionWenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng et al.ICCV 2025 · 2 citations
- GOLD: Graph Out-of-Distribution Detection via Implicit Adversarial Latent GenerationDanny Wang, Ruihong Qiu, Guangdong Bai, Zi HuangICLR 2025
- A Semi-supervised Learning Approach with Two Teachers to Improve Breakdown Identification in DialoguesQian Lin, Hwee Tou NgAAAI 2022 · 5 citations
- Likelihood Ratios and Generative Classifiers for Unsupervised Out-of-Domain Detection in Task Oriented DialogVarun Gangal, Abhinav Arora, Arash Einolghozati, Sonal GuptaAAAI 2020 · 59 citations
