AdaStreamLite: Environment-adaptive Streaming Speech Recognition on Mobile Devices
Yuheng Wei, Jie Xiong, Hui Liu, Yingtao Yu, Jiangtao Pan, Junzhao Du
摘要
Streaming speech recognition aims to transcribe speech to text in a streaming manner, providing real-time speech interaction for smartphone users. However, it is not trivial to develop a high-performance streaming speech recognition system purely running on mobile platforms, due to the complex real-world acoustic environments and the limited computational resources of smartphones. Most existing solutions lack the generalization to unseen environments and have difficulty to work with streaming speech. In this paper, we design AdaStreamLite, an environment-adaptive streaming speech recognition tool for smartphones. AdaStreamLite interacts with its surroundings to capture the characteristics of the current acoustic environment to improve the robustness against ambient noise in a lightweight manner. We design an environment representation extractor to model acoustic environments with compact feature vectors, and construct a representation lookup table to improve the generalization of AdaStreamLite to unseen environments. We train our system using large speech datasets publicly available covering different languages. We conduct experiments in a large range of real acoustic environments with different smartphones. The results show that AdaStreamLite outperforms the state-of-the-art methods in terms of recognition accuracy, computational resource consumption and robustness against unseen environments.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- UltraSpeech: Speech Enhancement by Interaction between Ultrasound and SpeechHan Ding, Yizhan Wang, Hao Li, Cui Zhao 等UbiComp 2022 · 被引用 32 次
- AdaSpring: Context-adaptive and Runtime-evolutionary Deep Model Compression for Mobile ApplicationsSicong Liu, Bin Guo, Ke Ma, Zhiwen Yu 等UbiComp 2021 · 被引用 30 次
- Sensing to Hear: Speech Enhancement for Mobile Devices Using Acoustic SignalsQian Zhang, Dong Wang, Run Zhao, Yinggang Yu 等UbiComp 2021 · 被引用 22 次
- Global Normalization for Streaming Speech Recognition in a Modular FrameworkEhsan Variani, Ke Wu, Michael D. Riley, David Rybach 等NeurIPS 2022 · 被引用 15 次
相关 Paper
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar 等ICML 2023 · 被引用 12 次
- AdaEnlight: Energy-aware Low-light Video Stream Enhancement on Mobile DevicesSicong Liu, Xiaochen Li, Zimu Zhou, Bin Guo 等UbiComp 2023 · 被引用 16 次
- Real-Time Neural Voice CamouflageMia Chiquier, Chengzhi Mao, Carl VondrickICLR 2022 · 被引用 9 次
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 被引用 54 次
- Audio Lottery: Speech Recognition Made Ultra-Lightweight, Noise-Robust, and TransferableShaojin Ding, Tianlong Chen, Zhangyang WangICLR 2022 · 被引用 34 次
