Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
Zekun Li, Zhiyu Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Dong, Adithya Sagar, Xifeng Yan, Paul A. Crook
Abstract
Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. However, their effectiveness in task-oriented dialogues (TOD), which requires not only response generation but also effective dialogue state tracking (DST) within specific tasks and domains, remains less satisfying. In this work, we propose a novel approach FNCTOD for solving DST with LLMs through function calling. This method improves zero-shot DST, allowing adaptation to diverse domains without extensive data collection or model tuning. Our experimental results demonstrate that our approach achieves exceptional performance with both modestly sized open-source and also proprietary LLMs: with in-context prompting it enables various 7B or 13B parameter models to surpass the previous state-of-the-art (SOTA) achieved by ChatGPT, and improves ChatGPT's performance beating the SOTA by 5.6% average joint goal accuracy (JGA). Individual model results for GPT-3.5 and GPT-4 are boosted by 4.8% and 14%, respectively. We also show that by fine-tuning on a small collection of diverse task-oriented dialogues, we can equip modestly sized models, specifically a 13B parameter LLaMA2-Chat model, with function-calling capabilities and DST performance comparable to ChatGPT while maintaining their chat capabilities. We have made the code publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f0b12be3-c6bb-4e72-8a84-38eacc6eafa3Cited by top-tier papers13
- Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMsReham Omar, Omij Mangukiya, Essam MansourSIGMOD 2025 · 9 citations
- GraphTool-Instruction: Revolutionizing Graph Reasoning in LLMs through Decomposed Subtask InstructionRongzheng Wang, Shuang Liang, Qizhi Chen, Jiasheng Zhang et al.KDD 2025 · 6 citations
- Zero-shot Cross-domain Dialogue State Tracking via Context-aware Auto-prompting and Instruction-following Contrastive DecodingXiaoyu Dong, Yujie Feng, Zexin Lu, Guangyuan Shi et al.EMNLP 2024 · 3 citations
- Unsupervised End-to-End Task-Oriented Dialogue with LLMs: The Power of the Noisy ChannelBrendan King, Jeffrey FlaniganEMNLP 2024 · 3 citations
- CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow AlignmentRadin Shayanfar, Chu Fei Luo, Rohan Bhambhoria, Samuel Dahan et al.ACL 2026 · 1 citation
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
Related papers
- Towards LLM-driven Dialogue State TrackingYujie Feng, Zexin Lu, Bo Liu, Liming Zhan et al.EMNLP 2023 · 25 citations
- Enhancing Dialogue State Tracking Models through LLM-backed User-Agents SimulationCheng Niu, Xingguang Wang, Xuxin Cheng, Juntong Song et al.ACL 2024
- Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPTXiaoshuai Song, Keqing He, Pei Wang, Guanting Dong et al.EMNLP 2023 · 3 citations
- DiSTRICT: Dialogue State Tracking with Retriever Driven In-Context TuningPraveen Venkateswaran, Evelyn Duesterwald, Vatche IsahagianEMNLP 2023 · 7 citations
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
