DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input
Jun Rekimoto
Abstract
Interactions based on automatic speech recognition (ASR) have become widely used, with speech input being increasingly utilized to create documents. However, as there is no easy way to distinguish between commands being issued and text required to be input in speech, misrecognitions are difficult to identify and correct, meaning that documents need to be manually edited and corrected. The input of symbols and commands is also challenging because these may be misrecognized as text letters. To address these problems, this study proposes a speech interaction method called DualVoice, by which commands can be input in a whispered voice and letters in a normal voice. The proposed method does not require any specialized hardware other than a regular microphone, enabling a complete hands-free interaction. The method can be used in a wide range of situations where speech recognition is already available, ranging from text input to mobile/wearable computing. Two neural networks were designed in this study, one for discriminating normal speech from whispered speech, and the second for recognizing whisper speech. A prototype of a text input system was then developed to show how normal and whispered voice can be used in speech text input. Other potential applications using DualVoice are also discussed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdcf4ab1-8316-4a1e-8f64-9cd491f6e9abCited by top-tier papers2
- WESPER: Zero-shot and Realtime Whisper to Normal Voice Conversion for Whisper-based Speech InteractionsJun RekimotoCHI 2023 · 24 citations
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational AgentsNikhil Sharma, Zheng Zhang, Daniel Lee, Namita Krishnan et al.CHI 2026 · 2 citations
Builds on2
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single MicrophoneYue Qin, Chun Yu, Zhaoheng Li, Mingyuan Zhong et al.CHI 2021 · 21 citations
Related papers
- NasoVoce: A Nose-Mounted Low-Audibility Speech Interface for Always-Available Speech InteractionJun Rekimoto, Yu Nishimura, Bojian YangCHI 2026 · 1 citation
- EarCommand: "Hearing" Your Silent Speech Commands In EarYincheng Jin, Yang Gao, Xuhai Xu, Seokmin Choi et al.UbiComp 2022 · 37 citations
- SoundLip: Enabling Word and Sentence-level Lip Interaction for Smart DevicesQian Zhang, Dong Wang, Run Zhao, Yinggang YuUbiComp 2021 · 34 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- EchoWhisper: Exploring an Acoustic-based Silent Speech Interface for Smartphone UsersYang Gao, Yincheng Jin, Jiyang Li, Seokmin Choi et al.UbiComp 2020 · 44 citations
