Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
Hyunjong Ok, Suho Yoo, Jaeho Lee
Abstract
Spoken dialogue systems powered by large language models have demonstrated remarkable abilities in understanding human speech and generating appropriate spoken responses. However, these systems struggle with end-turn detection (ETD) -- the ability to distinguish between user turn completion and hesitation. This limitation often leads to premature or delayed responses, disrupting the flow of spoken conversations. In this paper, we introduce the ETD Dataset, the first public dataset for end-turn detection. The ETD dataset consists of both synthetic speech data generated with text-to-speech models and real-world speech data collected from web sources. We also propose SpeculativeETD, a novel collaborative inference framework that balances efficiency and accuracy to improve real-time ETD in resource-constrained environments. Our approach jointly employs a lightweight GRU-based model, which rapidly detects the non-speaking units in real-time on local devices, and a high-performance Wav2vec-based model running on the server to make a more challenging classification of distinguishing turn ends from mere pauses. Experiments demonstrate that the proposed SpeculativeETD significantly improves ETD accuracy while keeping the required computations low. Datasets and code will be available after the review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34173e4d-1d3c-4ac7-9bee-f9c68ef8f0f6Builds on9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- Can Large Language Model Agents Simulate Human Trust Behavior?Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye et al.NeurIPS 2024 · 183 citations
Related papers
- Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual SignalsYuxin Lin, Yinglin Zheng, Ming Zeng, Wangzheng ShiACL 2025 · 5 citations
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- Aligning Spoken Dialogue Models from User InteractionsAnne Wu, Laurent Mazaré, Neil Zeghidour, Alexandre DéfossezICML 2025
- AV-Dialog: Spoken Dialogue Models with Audio-Visual InputTuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath GollakotaACL 2026 · 1 citation
- Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking DynamicsSiddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang et al.ICLR 2025
