JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
Jianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun, Ruihua Song, Liyun Ru
Abstract
Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information, including vision and audio, an effective benchmark must comprehensively cover three key aspects: (1) multi-modal dependency (i.e., questions that cannot be answered using vision or audio alone), (2) diverse audio information types (e.g., speech, sound events), and (3) varying scene spans. However, existing datasets fall short in one or more of these dimensions, limiting strict and comprehensive evaluation. To address this gap, we introduce JointAVBench, a novel benchmark with strict audio-video correlation, spanning five cognitive dimensions, four audio information types (speech, sound events, music, vocal traits), and three scene spans (single-, cross-, and full-scene). Given the high cost of manual annotation, we propose an automated pipeline that leverages state-of-the-art vision-LLMs, audio-LLMs, and general-purpose LLMs to synthesize questions and answers that strictly require joint audio-visual understanding. We evaluate leading vision-only, audio-only, and Omni-LLMs on our dataset. Results show that even the best-performing Omni-LLM achieves an average accuracy of only 65.3%, outperforming uni-modal baselines but revealing substantial room for improvement, especially in cross-scene reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d27c8fa4-1212-44e3-aae0-b9c653a9832cCited by top-tier papers2
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang et al.ICML 2026 · 5 citations
- PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio PerceptionTianwei Lan, Jiaqi Wu, Zeming Liu, Zhaoxin Fan et al.ACL 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- NExT-GPT: Any-to-Any Multimodal LLMShengqiong Wu, Hao Fei, Leigang Qu, Wei Ji et al.ICML 2024 · 786 citations
Related papers
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMsYaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu et al.ICML 2026
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu et al.ICLR 2026 · 53 citations
- AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language ModelsSung-Bin Kim, Oh Hyun-Bin, JungMok Lee, Arda Senocak et al.ICLR 2025
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang et al.ICML 2026 · 8 citations
