SSAST: Self-Supervised Audio Spectrogram Transformer
Yuan Gong, Cheng-I Lai, Yu-An Chung, James R. Glass
Abstract
Recently, neural networks based purely on self-attention, such as the Vision Transformer (ViT), have been shown to outperform deep learning models constructed with convolutional neural networks (CNNs) on various vision tasks, thus extending the success of Transformers, which were originally developed for language processing, to the vision domain. A recent study (Gong, Chung, and Glass 2021) showed that a similar methodology can also be applied to the audio domain. Specifically, the Audio Spectrogram Transformer (AST) achieves state-of-the-art results on various audio classification benchmarks. However, pure Transformer models tend to require more training data compared to CNNs, and the success of the AST relies on supervised pretraining that requires a large amount of labeled data and a complex training pipeline, thus limiting the practical usage of AST. This paper focuses on audio and speech classification, and aims to reduce the need for large amounts of labeled data for the AST by leveraging self-supervised learning using unlabeled data. Specifically, we propose to pretrain the AST model with joint discriminative and generative masked spectrogram patch modeling (MSPM) using unlabeled audio from AudioSet and Librispeech. We evaluate our pretrained models on both audio and speech classification tasks including audio event classification, keyword spotting, emotion recognition, and speaker identification. The proposed self-supervised framework significantly boosts AST performance on all tasks, with an average improvement of 60.9%, leading to similar or even better results than a supervised pretrained AST. To the best of our knowledge, it is the first patch-based selfsupervised learning framework in the audio and speech domain, and also the first self-supervised learning framework for AST. Code at https://github.com/YuanGongND/ssast .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c90002e1-e818-4668-8043-7e563aa6763fCited by top-tier papers35
- BEATs: Audio Pre-Training with Acoustic TokenizersSanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu et al.ICML 2023 · 568 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 352 citations
- MAViL: Masked Audio-Video LearnersPo-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali et al.NeurIPS 2023 · 95 citations
- TVLT: Textless Vision-Language TransformerZineng Tang, Jaemin Cho, Yixin Nie, Mohit BansalNeurIPS 2022 · 40 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- DTF-AT: Decoupled Time-Frequency Audio Transformer for Event ClassificationTony Alex, Sara Ahmed, Armin Mustafa, Muhammad Awais et al.AAAI 2024 · 8 citations
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie et al.ICML 2026 · 2 citations
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng et al.ACM MM 2020 · 93 citations
- Patch-level Representation Learning for Self-supervised Vision TransformersSukmin Yun, Hankook Lee, Jaehyung Kim, Jinwoo ShinCVPR 2022 · 52 citations
- Contrastive Audio-Visual Masked AutoencoderYuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath et al.ICLR 2023 · 17 citations
