MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
Ramaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
Abstract
The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more contextaware interactions. However, current benchmarks fall short in comprehensively evaluating how well these models generate context-aware responses, particularly when it comes to implicitly understanding fine-grained speech characteristics, such as pitch, emotion, timbre, and volume or the environmental acoustic context such as background sounds. Additionally, they inadequately assess the ability of models to align paralinguistic cues with complementary visual signals to inform their responses. To address these gaps, we introduce MULTIVOX, the first omni voice assistant benchmark designed to evaluate the ability of voice assistants to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. Specifically, MULTIVOX includes 1000 human-annotated and recorded speech dialogues that encompass diverse paralinguistic features and a range of visual cues such as images and videos. Our evaluation on 10 state-of-the-art models reveals that, although humans excel at these tasks, current open-source models consistently struggle to produce contextually grounded responses. 1 General Intelligence (AGI) (Bubeck et al., 2023; Morris et al., 2024) . While OLMs provide a wide range of applications (Xu et al., 2025) , one of their primary use cases is developing omni-modal voice assistants (OVA) (Huang et al., 2024) . Unlike traditional speech voice assistants that rely solely on speech instruction, OVAs powered by OLMs such as GPT-4o (OpenAI, 2024) and Qwen2.5 Omni (Xu et al., 2025), can understand speech dialogues and reason over multimodal inputs, including images and videos. Advancing the application of OLMs in voice assistants poses challenges not only in model development but also in constructing effective evaluation benchmarks. While existing OLM benchmarks like OmniBench (Li et al., 2024) incorpo-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ed98b2e-1866-4f95-9bb3-3f9e1bf40102Cited by top-tier papers2
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue et al.ACL 2026 · 17 citations
- Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadoxJiacheng Pang, Ashutosh Chaubey, Mohammad SoleymaniICML 2026 · 5 citations
Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMEliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar et al.ICLR 2024 · 95 citations
- Vision-Speech Models: Teaching Speech Models to Converse about ImagesAmélie Royer, Moritz Böhle, Laurent Mazaré, Neil Zeghidour et al.CVPR 2026 · 4 citations
- VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?Xize Cheng, Ruofan Hu, Xiaoda Yang, Jingyu Lu et al.ICLR 2025
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning BenchmarkS. Sakshi, Utkarsh Tyagi, Sonal Kumar, Ashish Seth et al.ICLR 2025
Related papers
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMsYaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu et al.ICML 2026
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning EvaluationJianghan Chao, Jianzhang Gao, Wenhui Tan, Yuchong Sun et al.ICLR 2026 · 16 citations
- Omni-RewardBench: Toward a Comprehensive Evaluation of Generative Reward Models Across ModalitiesChi-Min Chan, Yujin Zhou, Pengcheng Wen, Boqin Yin et al.ACL 2026
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning BenchmarkDingdong Wang, Junan Li, Jincenzi Wu, Dongchao Yang et al.ICLR 2026 · 143 citations
