Rethinking Task-Oriented Dialogue Systems: From Complex Modularity to Zero-Shot Autonomous Agent
Heng-Da Xu, Xian-Ling Mao, Puhai Yang, Fanshu Sun, Heyan Huang
摘要
Task-oriented dialogue (TOD) systems are predominantly designed to be composed of several functional modules (e.g. dialogue state tracker, dialogue policy, natural language generation) whether they are pipeline or end-to-end architectures. However, this modular design not only heavily relies on massive fully-annotated data, but also suffers from many intrinsic drawbacks, such as serious error accumulation, poor generalization ability, high customization cost, and low fault tolerance rate. In this paper, we rethink the architecture of the task-oriented dialogue systems and propose a novel fully zeroshot autonomous TOD agent, named AutoTOD, where all the delicate modules in traditional TOD systems are deprecated and all it needs is a general-purpose instruction-following language model (e.g. GPT-4). AutoTOD only leverages a simple instruction schema consisting of the description of tasks and external APIs, and can autonomously decide to what to do at each dialogue turn, including asking for information, calling APIs, summarizing API results, and correcting previous mistakes. Moreover, we propose a simulation-based evaluation framework to better validate the abilities of TOD models in real-life scenarios. Extensive experiments conducted on the MultiWOZ and SGD datasets show the superior task completion ability and flexible language skills of AutoTOD. 1 040 2020). Traditional TOD systems are mostly de-041 signed as a pipeline of several separate modules, 042 including natural language understanding, dialogue 043 state tracker, dialogue policy, and natural language 044 generation (Zhang et al., 2020). These modules are 045 trained separately and work one by one to generate 046 the dialogue response to the user (Su et al., 2022). 047 Later, end-to-end TOD systems emerged where the 048 separate modules are combined and built on a sin-049 gle pretrained language model (He et al., 2022a; 050 Yang et al., 2021). Thus the whole system can be 051 trained end-to-end with annotated task dialogues. 052 Examples of these two kinds of TOD systems are 053 shown in Figure 1 (a, b). Nevertheless, both the 054 pipeline and end-to-end models are essentially in 055 the same modular architecture. 056 127 2019). The results show the superior task comple-128 tion ability and fluent language skills of AutoTOD. 129 Furthermore, AutoTOD demonstrates great robust-130 ness when facing various dialogue scenarios. 131 2 Related Work 132 2.1 Task-Oriented Dialogue Systems 133 Task-oriented dialogue (TOD) systems have been 134 studied for decades. Traditional approaches are fun-135 damentally built in a pipeline architecture, consist-136 ing of components including natural language un-137 derstanding, dialogue state tracking, dialogue pol-138 icy learning, and natural language generation (Wu 139
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM AgentsYiming Du, Bingbing Wang, Yang He, Bin Liang 等AAAI 2026 · 被引用 2 次
- Reasoning Gets Harder for LLMs Inside A DialogueIvan Kartác, Mateusz Lango, Ondrej DusekACL 2026 · 被引用 2 次
- EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue SystemsZhengyi Zhao, Shubo Zhang, Yiming Du, Bin Liang 等ACL 2026 · 被引用 2 次
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuningYiming Du, Yifan Xiang, Bin Liang, Dahua Lin 等EMNLP 2025 · 被引用 1 次
- ReContraster: Making Your Posters Stand Out with Regional ContrastPeixuan Zhang, Zijian Jia, Ziqi Cai, Shuchen Weng 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper5
- GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy InjectionWanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu 等AAAI 2022 · 被引用 181 次
- The AI Doctor Is In: A Survey of Task-Oriented Dialogue Systems for Healthcare ApplicationsMina Valizadeh, Natalie PardeACL 2022 · 被引用 55 次
- Multi-Domain Dialogue Acts and Response Co-GenerationKai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan 等ACL 2020 · 被引用 46 次
- Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and GenerationWanwei He, Yinpei Dai, Min Yang, Jian Sun 等SIGIR 2022 · 被引用 41 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
相关 Paper
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta 等ACL 2022 · 被引用 218 次
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz 等NeurIPS 2020 · 被引用 590 次
- META-GUI: Towards Multi-modal Conversational Agents on Mobile GUILiangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai 等EMNLP 2022 · 被引用 11 次
- AnyTOD: A Programmable Task-Oriented Dialog SystemJeffrey Zhao, Yuan Cao, Raghav Gupta, Harrison Lee 等EMNLP 2023 · 被引用 4 次
- TOD-Flow: Modeling the Structure of Task-Oriented DialoguesSungryull Sohn, Yiwei Lyu, Anthony Z. Liu, Lajanugen Logeswaran 等EMNLP 2023 · 被引用 3 次
