Multimodal Contextualized Semantic Parsing from Speech
Jordan Voas, David Harwath, Raymond Mooney
Abstract
We introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts. SPICE goes beyond traditional semantic parsing by offering a structured, interpretable framework for dynamically updating an agent's knowledge with new information, mirroring the complexity of human communication. We develop the VG-SPICE dataset, crafted to challenge agents with visual scene graph construction from spoken conversational exchanges, highlighting speech and visual data integration. We also present the Audio-Vision Dialogue Scene Parser (AViD-SP) developed for use on VG-SPICE. These innovations aim to improve multimodal information processing and integration. Both the VG-SPICE dataset and the AViD-SP model are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4931ea5-bf34-49ba-ad63-e504c3e7ab57Builds on15
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 9 citations
- Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled TransformersShijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori et al.AAAI 2021 · 50 citations
- DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph RefinementShaoqing Lin, Chong Teng, Fei Li, Donghong Ji et al.EMNLP 2025
- VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsYuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li et al.ACL 2023 · 4 citations
- Exploring Contextual-Aware Representation and Linguistic-Diverse Expression for Visual DialogXiangpeng Li, Lianli Gao, Lei Zhao, Jingkuan SongACM MM 2021 · 3 citations
