M3Net: Efficient Time-Frequency Integration Network with Mirror Attention for Audio Classification on Edge
Xuanming Jiang, Baoyi An, Guoshuai Zhao, Xueming Qian
摘要
Audio classification plays a crucial role within fields such as human-machine interaction and intelligent robotics. However, high-performance audio classification systems typically demand significant computational and storage resources, posing substantial challenges when deploying to the resource-constrained edge devices with an urgent need for such capabilities. To achieve a new level of balance between model complexity and performance, we introduce a novel multi-view method for the separated time-frequency features extraction and utilization, which exists within the proposed Mini Mirror Multi-View Network (M3Net) in the form of the Mirror Attention mechanism. M3Net enables reversible spatial transformation of spectral features is capable of efficiently leverages robust local and global features in the time and frequency domains with low requirements for parameters. Experiments based on Mel-Spectrogram without data augmentation and pre-training indicate that M3Net can achieve classification accuracy over 97% on the UrbanSound8K and SpeechCommandsV2 datasets with only 0.03 million parameters. The contribution of each functional segment in M3Net is fully verified and explained in the ablation experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- SSAST: Self-Supervised Audio Spectrogram TransformerYuan Gong, Cheng-I Lai, Yu-An Chung, James R. GlassAAAI 2022 · 被引用 397 次
- Learning Delays in Spiking Neural Networks using Dilated Convolutions with Learnable SpacingsIlyass Hammouamri, Ismail Khalfaoui Hassani, Timothée MasquelierICLR 2024 · 被引用 105 次
- Sparse Modular Activation for Efficient Sequence ModelingLiliang Ren, Yang Liu, Shuohang Wang, Yichong Xu 等NeurIPS 2023 · 被引用 23 次
- CoLLAT: On Adding Fine-grained Audio Understanding to Language Models using Token-Level Locked-Language TuningDadallage A. R. Silva, Spencer Whitehead, Christopher T. Lengerich, Hugh LeatherNeurIPS 2023 · 被引用 11 次
相关 Paper
- RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationSamuel Pegg, Kai Li, Xiaolin HuICLR 2024 · 被引用 13 次
- IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationKai Li, Runxuan Yang, Fuchun Sun, Xiaolin HuICML 2024 · 被引用 28 次
- DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationTony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais 等AAAI 2024 · 被引用 8 次
- An efficient encoder-decoder architecture with top-down attention for speech separationKai Li, Runxuan Yang, Xiaolin HuICLR 2023 · 被引用 16 次
- Audio-Visual Glance Network for Efficient Video RecognitionMuhammad Adi Nugroho, Sangmin Woo, Sumin Lee, Changick KimICCV 2023 · 被引用 8 次
