Short-Term Memory Convolutions
Grzegorz Stefanski, Krzysztof Arendt, Pawel Daniluk, Bartlomiej Jasik, Artur Szumaczuk
摘要
The real-time processing of time series signals is a critical issue for many reallife applications. The idea of real-time processing is especially important in audio domain as the human perception of sound is sensitive to any kind of disturbance in perceived signals, especially the lag between auditory and visual modalities. The rise of deep learning (DL) models complicated the landscape of signal processing. Although they often have superior quality compared to standard DSP methods, this advantage is diminished by higher latency. In this work we propose novel method for minimization of inference time latency and memory consumption, called Short-Term Memory Convolution (STMC) and its transposed counterpart. The main advantage of STMC is the low latency comparable to long short-term memory (LSTM) networks. Furthermore, the training of STMC-based models is faster and more stable as the method is based solely on convolutional neural networks (CNNs). In this study we demonstrate an application of this solution to a U-Net model for a speech separation task and GhostNet model in acoustic scene classification (ASC) task. In case of speech separation we achieved a 5-fold reduction in inference time and a 2-fold reduction in latency without affecting the output quality. The inference time for ASC task was up to 4 times faster while preserving the original accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- Locality defeats the curse of dimensionality in convolutional teacher-student scenariosAlessandro Favero, Francesco Cagnetta, Matthieu WyartNeurIPS 2021 · 被引用 34 次
- GhostNet: More Features From Cheap OperationsKai Han, Yunhe Wang, Qi Tian, Jianyuan Guo 等CVPR 2020
- MoViNets: Mobile Video Networks for Efficient Video RecognitionDan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang 等CVPR 2021
相关 Paper
- Don't Think It Twice: Exploit Shift Invariance for Efficient Online Streaming Inference of CNNsChristodoulos Kechris, Jonathan Dan, José Miranda, David AtienzaAAAI 2025 · 被引用 1 次
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D ScenesAnton Ratnarajah, Zhenyu Tang, Rohith Aralikatti, Dinesh ManochaACM MM 2022 · 被引用 35 次
- msf-CNN: Patch-based Multi-Stage Fusion with Convolutional Neural Networks for TinyMLZhaolan Huang, Emmanuel BaccelliNeurIPS 2025 · 被引用 4 次
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi 等AAAI 2020 · 被引用 151 次
