PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs
Rahul Goel, Waleed Ammar, Aditya Gupta, Siddharth Vashishtha, Motoki Sano, Faiz Surani, Max Chang, HyunJeong Choe, David Greene, Chuan He, Rattima Nitisaroj, Anna Trukhina
摘要
Research interest in task-oriented dialogs has increased as systems such as Google Assistant, Alexa and Siri have become ubiquitous in everyday life. However, the impact of academic research in this area has been limited by the lack of datasets that realistically capture the wide array of user pain points. To enable research on some of the more challenging aspects of parsing realistic conversations, we introduce PRESTO 1 , a public dataset of over 550K contextual multilingual conversations between humans and virtual assistants. PRESTO contains a diverse array of challenges that occur in realworld NLU tasks such as disfluencies, codeswitching, and revisions. It is the only largescale human generated conversational parsing dataset that provides structured context such as a user's contacts and lists for each example. Our mT5 model-based baselines demonstrate that the conversational phenomena present in PRESTO are challenging to model, which is further pronounced in a low-resource setup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- How Do Language Models Speak Languages? A Case Study on Unintended Code-SwitchingYuxin Xiao, Zhen Huang, Wenxiao Wang, Yan Zhao 等ICML 2026
- In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMsBhavani Shankar, Preethi Jyothi, Pushpak BhattacharyyaACL 2024
它引用的顶会 Paper4
- End-to-End Slot Alignment and Recognition for Cross-Lingual NLUWeijia Xu, Batool Haider, Saab MansourEMNLP 2020 · 被引用 109 次
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie 等ACL 2023 · 被引用 88 次
- Conversational Semantic Parsing for Dialog State TrackingJianpeng Cheng, Devang Agrawal, Héctor Martínez Alonso, Shruti Bhargava 等EMNLP 2020 · 被引用 41 次
- Controllable Semantic Parsing via Retrieval AugmentationPanupong Pasupat, Yuan Zhang, Kelvin GuuEMNLP 2021 · 被引用 29 次
相关 Paper
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- Code-switched inspired losses for spoken dialog representationsPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 被引用 6 次
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong 等ACL 2026
- GupShup: Summarizing Open-Domain Code-Switched ConversationsLaiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi 等EMNLP 2021 · 被引用 13 次
- TIMEDIAL: Temporal Commonsense Reasoning in DialogLianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He 等ACL 2021
