LEAF: A Learnable Frontend for Audio Classification
Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi
Abstract
Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fundamental limitations of handmade representations. In this work we show that we can train a single learnable frontend that outperforms mel-filterbanks on a wide range of audio signals, including speech, music, audio events and animal sounds, providing a general-purpose learned frontend for audio classification. To do so, we introduce a new principled, lightweight, fully learnable architecture that can be used as a drop-in replacement of mel-filterbanks. Our system learns all operations of audio features extraction, from filtering to pooling, compression and normalization, and can be integrated into any neural network at a negligible parameter cost. We perform multi-task training on eight diverse audio classification tasks, and show consistent improvements of our model over mel-filterbanks and previous learnable alternatives. Moreover, our system outperforms the current state-of-the-art learnable frontend on Audioset, with orders of magnitude fewer parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e19e3df7-d0cc-4a80-a51a-8e97f2113603Cited by top-tier papers7
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin et al.AAAI 2024 · 29 citations
- In Search for a Generalizable Method for Source Free Domain AdaptationMalik Boudiaf, Tom Denton, Bart van Merrienboer, Vincent Dumoulin et al.ICML 2023 · 26 citations
- Learning Temporal Resolution in Spectrogram for Audio ClassificationHaohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang et al.AAAI 2024 · 15 citations
- ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Jianhua Tao et al.ACM MM 2025 · 5 citations
- Utilizing Speaker Profiles for Impersonation Audio DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Yong Ren et al.ACM MM 2024 · 3 citations
Related papers
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma et al.AAAI 2022 · 58 citations
- Deep Audio Priors Emerge From Harmonic Convolutional NetworksZhoutong Zhang, Yunyun Wang, Chuang Gan, Jiajun Wu et al.ICLR 2020 · 32 citations
- Learnable Group Transform For Time-SeriesRomain Cosentino, Behnaam AazhangICML 2020 · 14 citations
- FlowDec: A flow-based full-band general audio codec with high perceptual qualitySimon Welker, Matthew Le, Ricky T. Q. Chen, Wei-Ning Hsu et al.ICLR 2025
- Deep Edge Filter: Return of the Human-Crafted Layer in Deep LearningDongkwan Lee, JunHoo Lee, Nojun KwakNeurIPS 2025
