WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
Zhaojiang Lin, Yong Xu, Kai Sun, Jing Zheng, Yin Huang, Surya Teja Appini, Krish Narang, Renjie Tao, Ishan Kapil Jain, Siddhant Arora, Ruizhi Li, Yiteng Huang
Abstract
Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and noise, rapid micro-interactions, and the need to distinguish device-directed speech from background conversations. Existing benchmarks largely overlook these complexities, focusing instead on clean or generic conversational audio. To bridge this gap, we present WearVox, the first benchmark designed to rigorously evaluate voice assistants in realistic wearable scenarios. WearVox comprises 3,842 multi-channel, egocentric audio recordings collected via AI glasses across five diverse tasks including Search-Grounded QA, Closed-Book QA, Side-Talk Rejection, Tool Calling, and Speech Translation, spanning a wide range of indoor and outdoor environments and acoustic conditions. Each recording is accompanied by rich metadata, enabling nuanced analysis of model performance under real-world constraints. We benchmark leading proprietary and open-source speech Large Language Models (SLLMs) and find that most real-time SLLMs achieve accuracies on WearVox ranging from 29% to 59%, with substantial performance degradation on noisy outdoor audio, underscoring the difficulty and realism of the benchmark. Additionally, we conduct a case study with two new SLLMs that perform inference with single-channel and multi-channel audio, demonstrating that multi-channel audio inputs significantly enhance model robustness to environmental noise and improve discrimination between device-directed and background speech. Our results highlight the critical importance of spatial audio cues for context-aware voice assistants and establish WearVox as a comprehensive testbed for advancing wearable voice AI research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30671b2e-2132-4d12-a6ad-97f6a85a1448Builds on4
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu et al.ICML 2022 · 245 citations
- Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue AgentsBandhav Veluri, Benjamin N. Peloquin, Bokai Yu, Hongyu Gong et al.EMNLP 2024 · 8 citations
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji et al.ICML 2025
Related papers
- MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal InteractionsRamaneswaran Selvakumar, Ashish Seth, Nishit Anand, Utkarsh Tyagi et al.EMNLP 2025
- Benchmarking Egocentric Visual-Inertial SLAM at City ScaleAnusha Krishnan, Shaohui Liu, Paul-Edouard Sarlin, Oscar Gentilhomme et al.ICCV 2025 · 4 citations
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang et al.ICLR 2026 · 5 citations
- ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory PerceptionsBufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu et al.NeurIPS 2025 · 68 citations
- AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic FeaturesWeiye Xu, Zhang Jiang, Siqi Zheng, Xiyuxing Zhang et al.UbiComp 2026
