Towards Automated Error Discovery: A Study in Conversational AI
Dominic Petrak, Thy Thy Tran, Iryna Gurevych
Abstract
Although LLM-based conversational agents demonstrate strong fluency and coherence, they still produce undesirable behaviors (errors) that are challenging to prevent from reaching users during deployment. Recent research leverages large language models (LLMs) to detect errors and guide response-generation models toward improvement. However, current LLMs struggle to identify errors not explicitly specified in their instructions, such as those arising from updates to the response-generation model or shifts in user behavior. In this work, we introduce Automated Error Discovery, a framework for detecting and defining errors in conversational AI, and propose SEEED (Soft Clustering Extended Encoder-Based Error Detection), as an encoderbased approach to its implementation. We enhance the Soft Nearest Neighbor Loss by amplifying distance weighting for negative samples and introduce Label-Based Sample Ranking to select highly contrastive examples for better representation learning. SEEED outperforms adapted baselines-including GPT-4o and Phi-4-across multiple error-annotated dialogue datasets, improving the accuracy for detecting unknown errors by up to 8 points and demonstrating strong generalization to unknown intent detection. 1 1 We provide our code on GitHub: https://github.com/ UKPLab/emnlp2025-automatic-error-discovery. 1 1 2023; Mi et al., 2020; Roller et al., 2020) , these changes may lead to the emergence of new error types that the LLM might not recognize. In this work, we address the challenge of error detection in conversational AI. We introduce Automated Error Discovery, a framework for detecting and defining errors in dialogue, and propose SEEED (Soft Clustering Extended Encoder-Based Error Detection) as an approach to its implementation. Our contributions are as follows: • We introduce Automated Error Discovery, a framework for (1) detecting both known and unknown error types, and (2) generating definitions for newly discovered ones. • We propose SEEED, a novel approach that combines an open-source LLM with lightweight encoders for error detection. In contrast to prior work, SEEED employs soft clustering in the classification step, enabling more contextually coherent groupings. • We introduce Label-Based Sample Ranking, a novel sampling strategy for contrastive learning that selects highly contrastive examples based on the error they represent to improve representation learning. • We enhance the Soft Nearest Neighbor Loss (Frosst et al., 2019) by introducing a margin parameter to amplify the effect of distance weighting for negative samples. Oh, really? Yes, Rome is an impressive city. I also just came back from summer vacation. I did a lot of surfing! ? Summary Encoder I just came back from summer vacation. I've been to Rome. It's such a lovely city! Awesome, that sounds fun! Where did you go?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d859899-a159-468f-b3fb-9ffd9f6c0568Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
Related papers
- Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPTXiaoshuai Song, Keqing He, Pei Wang, Guanting Dong et al.EMNLP 2023 · 3 citations
- Watch the Neighbors: A Unified K-Nearest Neighbor Contrastive Learning Framework for OOD Intent DiscoveryYutao Mou, Keqing He, Pei Wang, Yanan Wu et al.EMNLP 2022 · 9 citations
- Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem SolutionsHang Li, Tianlong Xu, Kaiqi Yang, Yucheng Chu et al.ACL 2025
- Trial and Error: Exploration-Based Trajectory Optimization of LLM AgentsYifan Song, Da Yin, Xiang Yue, Jie Huang et al.ACL 2024
- Aegis: Automated Error Generation and Attribution for Multi-Agent SystemsFanqi Kong, Ruijie Zhang, Huaxiao Yin, Guibin Zhang et al.ICLR 2026 · 16 citations
