BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
Lukas Rauch, Raphael Schwinger, Moritz Wirth, René Heinrich, Denis Huseljic, Marek Herde, Jonas Lange, Stefan Kahl, Bernhard Sick, Sven Tomforde, Christoph Scholz
Abstract
Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a universaldomain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource. Therefore, we introduce BirdSet, a largescale benchmark dataset for audio classification focusing on avian bioacoustics. BirdSet surpasses AudioSet with over 6,800 recording hours (↑ 17%) from nearly 10,000 classes (↑ 18×) for training and more than 400 hours (↑ 7×) across eight strongly labeled evaluation datasets. It serves as a versatile resource for use cases such as multi-label classification, covariate shift, or self-supervised learning. We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification. We host our dataset on Hugging Face for easy accessibility and offer an extensive codebase to reproduce our results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 823b4009-e7f5-4eec-9072-4fa540bd9143Cited by top-tier papers8
- AVEX: What Matters for Animal Vocalization EncodingMarius Miron, David Robinson, Milad Alizadeh, Ellen Gilsenan-McMahon et al.ICLR 2026 · 9 citations
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio ClassificationLukas Rauch, René Heinrich, Houtan Ghaffari, Lukas Miklautz et al.ICLR 2026 · 7 citations
- WhAM: Towards A Translative Model of Sperm Whale VocalizationOrr Paradise, Liangyuan Chen, Pranav Muralikrishnan, Hugo Flores García et al.NeurIPS 2025 · 5 citations
- SounDiT: Geo-Contextual Soundscape-to-Landscape GenerationJunbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang et al.CVPR 2026 · 3 citations
- NatureLM-audio: an Audio-Language Foundation Model for BioacousticsDavid Robinson, Marius Miron, Masato Hagiwara, Olivier PietquinICLR 2025
Builds on4
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
Related papers
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma et al.AAAI 2022 · 58 citations
- A Sequential Self Teaching Approach for Improving Generalization in Sound Event RecognitionAnurag Kumar, Vamsi K. IthapuICML 2020 · 35 citations
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 224 citations
- In Search for a Generalizable Method for Source Free Domain AdaptationMalik Boudiaf, Tom Denton, Bart van Merrienboer, Vincent Dumoulin et al.ICML 2023 · 26 citations
- LEAF: A Learnable Frontend for Audio ClassificationNeil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco TagliasacchiICLR 2021 · 181 citations
