SLURP: A Spoken Language Understanding Resource Package
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena Rieser
Abstract
Spoken Language Understanding infers semantic meaning directly from audio data, and thus promises to reduce error propagation and misunderstandings in end-user applications. However, publicly available SLU resources are limited. In this paper, we release SLURP, a new SLU package containing the following: (1) A new challenging dataset in English spanning 18 domains, which is substantially bigger and linguistically more diverse than existing datasets; (2) Competitive baselines based on state-of-the-art NLU and ASR systems; (3) A new transparent metric for entity labelling which enables a detailed error analysis for identifying potential areas of improvement. SLURP is available at https: //github.com/pswietojanski/slurp
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8e59d565-9eca-44ea-ab44-81bcbcf0bea7Cited by top-tier papers37
- SALMONN: Towards Generic Hearing Abilities for Large Language ModelsChangli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen et al.ICLR 2024 · 557 citations
- Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and UnderstandingYifan Peng, Siddharth Dalmia, Ian R. Lane, Shinji WatanabeICML 2022 · 203 citations
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
- Textually Pretrained Speech Language ModelsMichael Hassid, Tal Remez, Tu Anh Nguyen, Itai Gat et al.NeurIPS 2023 · 117 citations
- Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language ModelZhen Ye, Peiwen Sun, Jiahe Lei, Hongzhan Lin et al.AAAI 2025 · 89 citations
Related papers
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding TasksSuwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad et al.ACL 2023 · 21 citations
- Text Is No More Enough! A Benchmark for Profile-Based Spoken Language UnderstandingXiao Xu, Libo Qin, Kaiji Chen, Guoxing Wu et al.AAAI 2022 · 9 citations
- ECLM: Entity Level Language Model for Spoken Language Understanding with Chain of IntentShangjian Yin, Peijie Huang, Jiatian Chen, Haojing Huang et al.ACL 2025 · 12 citations
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingYingmei Guo, Linjun Shou, Jian Pei, Ming Gong et al.EMNLP 2021 · 2 citations
