PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs
Rahul Goel, Waleed Ammar, Aditya Gupta, Siddharth Vashishtha, Motoki Sano, Faiz Surani, Max Chang, HyunJeong Choe, David Greene, Chuan He, Rattima Nitisaroj, Anna Trukhina
Abstract
Research interest in task-oriented dialogs has increased as systems such as Google Assistant, Alexa and Siri have become ubiquitous in everyday life. However, the impact of academic research in this area has been limited by the lack of datasets that realistically capture the wide array of user pain points. To enable research on some of the more challenging aspects of parsing realistic conversations, we introduce PRESTO 1 , a public dataset of over 550K contextual multilingual conversations between humans and virtual assistants. PRESTO contains a diverse array of challenges that occur in realworld NLU tasks such as disfluencies, codeswitching, and revisions. It is the only largescale human generated conversational parsing dataset that provides structured context such as a user's contacts and lists for each example. Our mT5 model-based baselines demonstrate that the conversational phenomena present in PRESTO are challenging to model, which is further pronounced in a low-resource setup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8467492-86c8-4855-bf7e-d31e40a08481Cited by top-tier papers2
- How Do Language Models Speak Languages? A Case Study on Unintended Code-SwitchingYuxin Xiao, Zhen Huang, Wenxiao Wang, Yan Zhao et al.ICML 2026
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
Builds on4
- End-to-End Slot Alignment and Recognition for Cross-Lingual NLUWeijia Xu, Batool Haider, Saab MansourEMNLP 2020 · 109 citations
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- Conversational Semantic Parsing for Dialog State TrackingJianpeng Cheng, Devang Agrawal, Héctor Martínez Alonso, Shruti Bhargava et al.EMNLP 2020 · 41 citations
- Controllable Semantic Parsing via Retrieval AugmentationPanupong Pasupat, Yuan Zhang, Kelvin GuuEMNLP 2021 · 29 citations
Related papers
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- Code-switched inspired losses for spoken dialog representationsPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 6 citations
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- GupShup: Summarizing Open-Domain Code-Switched ConversationsLaiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi et al.EMNLP 2021 · 13 citations
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He et al.ACL 2021
