BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
Lukas Rauch, Raphael Schwinger, Moritz Wirth, René Heinrich, Denis Huseljic, Marek Herde, Jonas Lange, Stefan Kahl, Bernhard Sick, Sven Tomforde, Christoph Scholz
摘要
Deep learning (DL) has greatly advanced audio classification, yet the field is limited by the scarcity of large-scale benchmark datasets that have propelled progress in other domains. While AudioSet is a pivotal step to bridge this gap as a universaldomain dataset, its restricted accessibility and limited range of evaluation use cases challenge its role as the sole resource. Therefore, we introduce BirdSet, a largescale benchmark dataset for audio classification focusing on avian bioacoustics. BirdSet surpasses AudioSet with over 6,800 recording hours (↑ 17%) from nearly 10,000 classes (↑ 18×) for training and more than 400 hours (↑ 7×) across eight strongly labeled evaluation datasets. It serves as a versatile resource for use cases such as multi-label classification, covariate shift, or self-supervised learning. We benchmark six well-known DL models in multi-label classification across three distinct training scenarios and outline further evaluation use cases in audio classification. We host our dataset on Hugging Face for easy accessibility and offer an extensive codebase to reproduce our results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- AVEX: What Matters for Animal Vocalization EncodingMarius Miron, David Robinson, Milad Alizadeh, Ellen Gilsenan-McMahon 等ICLR 2026 · 被引用 9 次
- Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio ClassificationLukas Rauch, René Heinrich, Houtan Ghaffari, Lukas Miklautz 等ICLR 2026 · 被引用 7 次
- WhAM: Towards A Translative Model of Sperm Whale VocalizationOrr Paradise, Liangyuan Chen, Pranav Muralikrishnan, Hugo Flores García 等NeurIPS 2025 · 被引用 5 次
- SounDiT: Geo-Contextual Soundscape-to-Landscape GenerationJunbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang 等CVPR 2026 · 被引用 3 次
- NatureLM-audio: an Audio-Language Foundation Model for BioacousticsDavid Robinson, Marius Miron, Masato Hagiwara, Olivier PietquinICLR 2025
它引用的顶会 Paper4
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu 等ICML 2023 · 被引用 568 次
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski 等NeurIPS 2022 · 被引用 524 次
相关 Paper
- Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataKe Chen, Xingjian Du, Bilei Zhu, Zejun Ma 等AAAI 2022 · 被引用 58 次
- A Sequential Self Teaching Approach for Improving Generalization in Sound Event RecognitionAnurag Kumar, Vamsi K. IthapuICML 2020 · 被引用 35 次
- Co-Separating Sounds of Visual ObjectsRuohan Gao, Kristen GraumanICCV 2019 · 被引用 224 次
- In Search for a Generalizable Method for Source Free Domain AdaptationMalik Boudiaf, Tom Denton, Bart van Merrienboer, Vincent Dumoulin 等ICML 2023 · 被引用 26 次
- LEAF: A Learnable Frontend for Audio ClassificationNeil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco TagliasacchiICLR 2021 · 被引用 181 次
