BREAK: Breaking the Dialogue State Tracking Barrier with Beam Search and Re-ranking
Seungpil Won, Heeyoung Kwak, Joongbo Shin, Janghoon Han, Kyomin Jung
Abstract
Despite the recent advances in dialogue state tracking (DST), the joint goal accuracy (JGA) of the existing methods on MultiWOZ 2.1 still remains merely 60%. In our preliminary error analysis, we find that beam search produces a pool of candidates that is likely to include the correct dialogue state. Motivated by this observation, we introduce a novel framework, called BREAK (Beam search and RE-rAnKing), that achieves outstanding performance on DST. Our proposed method performs DST in two stages: (i) generating k-best dialogue state candidates with beam search and (ii) re-ranking the candidates to select the correct dialogue state. This simple yet powerful framework shows state-of-the-art performance on all versions of MultiWOZ and M2M datasets. Most notably, we push the joint goal accuracy to 80-90% on MultiWOZ 2.1-2.4, which is an improvement of 23.6%, 26.3%, 21.7%, and 10.8% over the previous best-performing models, respectively. The data and code will be available at https://github.com/tony-won/DST -BREAK .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4945b62f-16ec-4772-84eb-d5993ae3eb2aCited by top-tier papers3
- Self-Calibrated Listwise Reranking with Large Language ModelsRuiyang Ren, Yuhao Wang, Kun Zhou, Wayne Xin Zhao et al.WWW 2025 · 12 citations
- Compress-then-Rank: Faster and Better Listwise Reranking with Large Language Models via Ranking-Aware Passage CompressionZhewei Zhi, Yingyi Zhang, Yizhen Jing, Xianneng Li et al.AAAI 2026 · 1 citation
- DeAL: Decoding-time Alignment for Large Language ModelsJames Y. Huang, Sailik Sengupta, Daniele Bonadiman, Yi-An Lai et al.ACL 2025
Builds on11
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz et al.NeurIPS 2020 · 590 citations
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
- Efficient Dialogue State Tracking by Selectively Overwriting MemorySungdong Kim, Sohee Yang, Gyuwan Kim, Sang-Woo LeeACL 2020 · 189 citations
- MinTL: Minimalist Transfer Learning for Task-Oriented Dialogue SystemsZhaojiang Lin, Andrea Madotto, Genta Indra Winata, Pascale FungEMNLP 2020 · 138 citations
Related papers
- Correctable-DST: Mitigating Historical Context Mismatch between Training and Inference for Improved Dialogue State TrackingHongyan Xie, Haoxiang Su, Shuangyong Song, Hao Huang et al.EMNLP 2022 · 10 citations
- Non-Autoregressive Dialog State TrackingHung Le, Richard Socher, Steven C. H. HoiICLR 2020 · 54 citations
- Dual Slot Selector via Local Reliability Verification for Dialogue State TrackingJinyu Guo, Kai Shuang, Jijie Li, Zihan WangACL 2021
- Multi-domain Dialogue State Tracking with Recursive InferenceLizi Liao, Tongyao Zhu, Le Hong Long, Tat-Seng ChuaWWW 2021 · 11 citations
- MetaASSIST: Robust Dialogue State Tracking with Meta LearningFanghua Ye, Xi Wang, Jie Huang, Shenghui Li et al.EMNLP 2022 · 10 citations
