Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT
Xiaoshuai Song, Keqing He, Pei Wang, Guanting Dong, Yutao Mou, Jingang Wang, Yunsen Xian, Xunliang Cai, Weiran Xu
Abstract
The tasks of out-of-domain (OOD) intent discovery and generalized intent discovery (GID) aim to extend a closed intent classifier to openworld intent sets, which is crucial to taskoriented dialogue (TOD) systems. Previous methods address them by fine-tuning discriminative models. Recently, although some studies have been exploring the application of large language models (LLMs) represented by ChatGPT to various downstream tasks, it is still unclear for the ability of ChatGPT to discover and incrementally extent OOD intents. In this paper, we comprehensively evaluate ChatGPT on OOD intent discovery and GID, and then outline the strengths and weaknesses of ChatGPT. Overall, ChatGPT exhibits consistent advantages under zero-shot settings, but is still at a disadvantage compared to fine-tuned models. More deeply, through a series of analytical experiments, we summarize and discuss the challenges faced by LLMs including clustering, domain-specific understanding, and cross-domain in-context learning scenarios. Finally, we provide empirical guidance for future directions to address these challenges. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 778342bd-1bd2-47ca-8a48-c05d52a8df59Cited by top-tier papers13
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu et al.ACL 2025 · 236 citations
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li et al.ACL 2024 · 39 citations
- Thought-Augmented Planning for LLM-Powered Interactive Recommender AgentHaocheng Yu, Yaxiong Wu, Hao Wang, Wei Guo et al.KDD 2026 · 16 citations
- Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language ModelsChao Xue, Yao Wang, Mengqiao Liu, Di Liang et al.ACL 2026 · 5 citations
- MuggleMath: Assessing the Impact of Query and Response Augmentation on Math ReasoningChengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong et al.ACL 2024 · 4 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 176 citations
- Discovering New Intents with Deep Aligned ClusteringHanlei Zhang, Hua Xu, Ting-En Lin, Rui LyuAAAI 2021 · 138 citations
- Discovering New Intents via Constrained Deep Adaptive Clustering with Cluster RefinementTing-En Lin, Hua Xu, Hanlei ZhangAAAI 2020 · 127 citations
Related papers
- Towards LLM-driven Dialogue State TrackingYujie Feng, Zexin Lu, Bo Liu, Liming Zhan et al.EMNLP 2023 · 25 citations
- Large Language Models as Zero-shot Dialogue State Tracker through Function CallingZekun Li, Zhiyu Chen, Mike Ross, Patrick Huber et al.ACL 2024 · 9 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang et al.AAAI 2024 · 57 citations
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang et al.AAAI 2024 · 30 citations
