SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
Mingfei Chen, Zijun Cui, Xiulong Liu, Jinlin Xiang, Yang Zheng, Jingyuan Li, Eli Shlizerman
Abstract
3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D scenes. We introduce SAVVY-Bench, the first benchmark for 3D spatial reasoning in dynamic scenes with synchronized spatial audio. SAVVY-Bench is comprised of thousands of carefully curated question-answer pairs probing both directional and distance relationships involving static and moving objects, and requires fine-grained temporal grounding, consistent 3D localization, and multimodal annotation. To tackle this challenge, we propose SAVVY, a novel trainingfree reasoning pipeline that consists of two stages: (i) Egocentric Spatial Tracks Estimation, which leverages AV-LLMs as well as other audio-visual methods to track the trajectories of key objects related to the query using both visual and spatial audio cues, and (ii) Dynamic Global Map Construction, which aggregates multimodal queried object trajectories and converts them into a unified global dynamic map. Using the constructed map, a final QA answer is obtained through a coordinate transformation that aligns the global map with the queried viewpoint. Empirical evaluation demonstrates that SAVVY substantially enhances performance of stateof-the-art AV-LLMs, setting a new standard and stage for approaching dynamic 3D spatial reasoning in AV-LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e581dbb-5f24-44a6-8b93-d757afdeded6Cited by top-tier papers4
- EgoSound: Benchmarking Sound Understanding in Egocentric VideosBingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun et al.CVPR 2026 · 7 citations
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 4 citations
- R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal SpaceTin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric SaxCVPR 2026 · 2 citations
- JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical EnvironmentsZhan Liu, Changli Tang, Yuxin WANG, Zhiyuan Zhu et al.ICML 2026
Builds on30
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao et al.ICLR 2024 · 1,170 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
Related papers
- 3DSRBENCH: A Comprehensive 3D Spatial Reasoning BenchmarkWufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou et al.ICCV 2025 · 15 citations
- Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View ScenesMohsen Gholami, Ahmad Rezaei, Zhou Weimin, Sitong Mao et al.ICLR 2026 · 67 citations
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan et al.CVPR 2026 · 2 citations
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline MatchingHao Zhong, Muzhi Zhu, Shenyan Zeng, Anzhou Li et al.CVPR 2026 · 1 citation
- Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D WorldYuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu et al.CVPR 2026 · 15 citations
