Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
Di Wu, Liting Jiang, Ruiyu Fang, Bianjing, Hongyan Xie, Haoxiang Su, Hao Huang, Zhongjiang He, Shuangyong Song, Xuelong Li
摘要
Spoken Language Understanding (SLU) consists of two sub-tasks: intent detection (ID) and slot filling (SF). Given its broad range of real-world applications, enhancing SLU for practical deployment is increasingly critical. Profile-based SLU addresses ambiguous user utterances by incorporating context awareness (CA), user profiles (UP), and knowledge graphs (KG) to support disambiguation, thereby advancing SLU research toward real-world applicability. However, existing SLU datasets still fall short in representing real-world scenarios. Specifically, (1) CA uses one-hot vectors for representation, which is overly idealized, and (2) models typically focuses solely on predicting intents and slot labels, neglecting the reasoning process that could enhance performance and interpretability. To overcome these limitations, we introduce VRSLU, a novel SLU dataset that integrates both Visual images and explicit Reasoning. For over-idealized CA, we use GPT-4o and FLUX.1-dev to generate images reflecting users’ environments and statuses, followed by human verification to ensure quality. For reasoning, GPT-4o is employed to generate explanations for predicted labels, which are then refined by human annotators to ensure accuracy and coherence. Additionally, we propose an instructional template, LR-Instruct, which first predicts labels and then generates corresponding reasoning. This two-step approach helps mitigate the influence of reasoning bias on label prediction. Experimental results confirm the effectiveness of incorporating visual information and highlight the promise of explicit reasoning in advancing SLU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SLURP: A Spoken Language Understanding Resource PackageEmanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena RieserEMNLP 2020 · 被引用 129 次
- End-to-End Slot Alignment and Recognition for Cross-Lingual NLUWeijia Xu, Batool Haider, Saab MansourEMNLP 2020 · 被引用 109 次
- DC-Instruct: An Effective Framework for Generative Multi-intent Spoken Language UnderstandingBowen Xing, Lizi Liao, Minlie Huang, Ivor W. TsangEMNLP 2024 · 被引用 9 次
- Text Is No More Enough! A Benchmark for Profile-Based Spoken Language UnderstandingXiao Xu, Libo Qin, Kaiji Chen, Guoxing Wu 等AAAI 2022 · 被引用 9 次
相关 Paper
- Uni-MIS: United Multiple Intent Spoken Language Understanding via Multi-View Intent-Slot InteractionShangjian Yin, Peijie Huang, Yuhong XuAAAI 2024 · 被引用 10 次
- Predicting Implicit Arguments in Procedural Video InstructionsAnil Batra, Laura Sevilla-Lara, Marcus Rohrbach, Frank KellerACL 2025
- ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal ModelsRohan Wadhawan, Hritik Bansal, Kai-Wei Chang, Nanyun PengICML 2024 · 被引用 20 次
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung 等ICCV 2025 · 被引用 1 次
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan 等NeurIPS 2025 · 被引用 18 次
