Learning Temporal Resolution in Spectrogram for Audio Classification
Haohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang, Mark D. Plumbley
Abstract
The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time Fourier Transform (STFT). Previous works generally assume the hop size should be a constant value (e.g., 10 ms). However, a fixed temporal resolution is not always optimal for different types of sound. The temporal resolution affects not only classification accuracy but also computational cost. This paper proposes a novel method, DiffRes, that enables differentiable temporal resolution modeling for audio classification. Given a spectrogram calculated with a fixed hop size, DiffRes merges non-essential time frames while preserving important frames. DiffRes acts as a "drop-in" module between an audio spectrogram and a classifier and can be jointly optimized with the classification task. We evaluate DiffRes on five audio classification tasks, using mel-spectrograms as the acoustic features, followed by off-the-shelf classifier backbones. Compared with previous methods using the fixed temporal resolution, the DiffRes-based method can achieve the equivalent or better classification accuracy with at least 25% computational cost reduction. We further show that DiffRes can improve classification accuracy by increasing the temporal resolution of input acoustic features, without adding to the computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 888c9b2d-632c-48d5-a3d3-a7fd0e978a1bCited by top-tier papers1
Ask how each one uses itBuilds on4
- SSAST: Self-Supervised Audio Spectrogram TransformerYuan Gong, Cheng-I Lai, Yu-An Chung, James R. GlassAAAI 2022 · 397 citations
- LEAF: A Learnable Frontend for Audio ClassificationNeil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco TagliasacchiICLR 2021 · 181 citations
- Learning Fast Sample Re-weighting Without Reward DataZizhao Zhang, Tomas PfisterICCV 2021 · 109 citations
- Learning Strides in Convolutional Neural NetworksRachid Riad, Olivier Teboul, David Grangier, Neil ZeghidourICLR 2022 · 54 citations
Related papers
- SCRAPL: Scattering Transform with Random Paths for Machine LearningChristopher Mitcheltree, Vincent Lostanlen, Emmanouil Benetos, Mathieu LagrangeICLR 2026
- 3D CNNs With Adaptive Temporal Feature ResolutionsMohsen Fayyaz, Emad Bahrami Rad, Ali Diba, Mehdi Noroozi et al.CVPR 2021
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 467 citations
- Diff-V2M: A Hierarchical Conditional Diffusion Model with Explicit Rhythmic Modeling for Video-to-Music GenerationShulei Ji, Zihao Wang, Jiaxing Yu, Xiangyuan Yang et al.AAAI 2026
- FASTER Recurrent Networks for Efficient Video ClassificationLinchao Zhu, Du Tran, Laura Sevilla-Lara, Yi Yang et al.AAAI 2020 · 60 citations
