LEAF: A Learnable Frontend for Audio Classification
Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi
摘要
Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental limitations of handmade representations. In this work we show that we can train a single learnable frontend that outperforms mel-filterbanks on a wide range of audio signals, including speech, music, audio events and animal sounds, providing a general-purpose learned frontend for audio classification. To do so, we introduce a new principled, lightweight, fully learnable architecture that can be used as a drop-in replacement of mel-filterbanks. Our system learns all operations of audio features extraction, from filtering to pooling, compression and normalization, and can be integrated into any neural network at a negligible parameter cost. We perform multi-task training on eight diverse audio classification tasks, and show consistent improvements of our model over mel-filterbanks and previous learnable alternatives. Moreover, our system outperforms the current state-of-the-art learnable frontend on Audioset, with orders of magnitude fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 等AAAI 2024 · 被引用 29 次
- In Search for a Generalizable Method for Source Free Domain AdaptationMalik Boudiaf, Tom Denton, Bart van Merrienboer, Vincent Dumoulin 等ICML 2023 · 被引用 26 次
- Learning Temporal Resolution in Spectrogram for Audio ClassificationHaohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang 等AAAI 2024 · 被引用 15 次
- ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao 等ACM MM 2025 · 被引用 5 次
- Utilizing Speaker Profiles for Impersonation Audio DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Yong Ren 等ACM MM 2024 · 被引用 3 次
相关 Paper
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
- Deep Audio Priors Emerge From Harmonic Convolutional NetworksZhoutong Zhang, Yunyun Wang, Chuang Gan, Jiajun Wu 等ICLR 2020 · 被引用 32 次
- Learnable Group Transform For Time-SeriesRomain Cosentino, Behnaam AazhangICML 2020 · 被引用 14 次
- FlowDec: A flow-based full-band general audio codec with high perceptual qualitySimon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu 等ICLR 2025
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep LearningDongkwan Lee, JunHoo Lee, Nojun KwakNeurIPS 2025
