EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
Ashish Seth, Ramaneswaran Selvakumar, S. Sakshi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
Abstract
In this paper, we present EH-MAM (Easyto-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and adaptive masking strategy. Specifically, during SSL training, we progressively introduce harder regions to the model for reconstruction. Our approach automatically selects hard regions and is built on the observation that the reconstruction loss of individual frames in MAM can provide natural signals to judge the difficulty of solving the MAM pre-text task for that frame. To identify these hard regions, we employ a teacher model that first predicts the frame-wise losses and then decides which frames to mask. By learning to create challenging problems, such as identifying harder frames and solving them simultaneously, the model is able to learn more effective representations and thereby acquire a more comprehensive understanding of the speech. Quantitatively, EH-MAM outperforms several state-ofthe-art baselines across various low-resource speech recognition and SUPERB benchmarks by 5%-10%. Additionally, we conduct a thorough analysis to show that the regions masked by EH-MAM effectively capture useful context across speech frames 1 . * Equal contribution, † Equal Ideation 1 Code: https://github.com/cs20s030/ehmam.git Random Masking Masked Frame Reconstruction Model Pre-defined masking strategy Model (Teacher) Masked Frame Reconstruction Model (Student) Selective Masking with Teacher (b) EH-MAM Pre-training (a) Traditional MAM Pre-training
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and LanguageAlexei Baevski, Arun Babu, Wei-Ning Hsu, Michael AuliICML 2023 · 137 citations
- RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-EncoderShitao Xiao, Zheng Liu, Yingxia Shao, Zhao CaoEMNLP 2022 · 63 citations
Related papers
- ExLM: Rethinking the Impact of [MASK] Tokens in Masked Language ModelsKangjie Zheng, Junwei Yang, Siyue Liang, Bin Feng et al.ICML 2025
- Self-supervised Masked Graph Autoencoder via Structure-aware CurriculumHaoyang Li, Xin Wang, Zeyang Zhang, Zongyuan Wu et al.ICML 2025
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim et al.EMNLP 2022 · 8 citations
- Adversarial Masking for Self-Supervised LearningYuge Shi, N. Siddharth, Philip H. S. Torr, Adam R. KosiorekICML 2022 · 110 citations
- Masked Autoencoders that ListenPo-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski et al.NeurIPS 2022 · 524 citations
