Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
Dongwook Lee, Eunwoo Song, Che Hyun Lee, Heeseung Kim, Sungroh Yoon
Abstract
While recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interruptions (TPI) from the primary user's ongoing flow, leaving them vulnerable to contextual failures. To bridge this gap, we introduce TPI-Train, a dataset of 88K instances designed with speaker-aware hard negatives to enforce acoustic cue prioritization for interruption handling, and TPI-Bench, a comprehensive evaluation framework designed to rigorously measure the interruption-handling strategy and precise speaker discrimination in deceptive contexts. Experiments demonstrate that our dataset design mitigates semantic shortcut learning-a critical pitfall where models exploit semantic context while neglecting acoustic signals essential for discerning speaker changes. We believe our work establishes a foundational resource for overcoming text-dominated unimodal reliance in SLMs, paving the way for more robust multi-party spoken interaction. The code for the framework is publicly available at https://tpi-va.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7b0e671-896a-429e-80d4-37e174e7b93bBuilds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Investigating Users' Preferences and Expectations for Always-Listening Voice AssistantsMadiha Tabassum, Tomasz Kosinski, Alisa Frik, Nathan Malkin et al.UbiComp 2020 · 74 citations
- OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen et al.ACL 2025 · 51 citations
- CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modallyDarina Koishigarina, Arnas Uselis, Seong Joon OhICLR 2026 · 33 citations
- A Mixed-Methods Approach to Understanding User Trust after Voice Assistant FailuresAmanda Baughan, Xuezhi Wang, Ariel Liu, Allison Mercurio et al.CHI 2023 · 31 citations
Related papers
- SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, Yohan JoACL 2026
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex ConversationsChengqian Ma, Wei Tao, Steven Y. GuoEMNLP 2025 · 7 citations
- VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language ModelsYuxiang Wang, HongYu Liu, Dekun Chen, Xueyao Zhang et al.ICLR 2026 · 5 citations
- SLURP: A Spoken Language Understanding Resource PackageEmanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena RieserEMNLP 2020 · 129 citations
