ReconVAT: A Semi-Supervised Automatic Music Transcription Framework for Low-Resource Real-World Data
Kin Wai Cheuk, Dorien Herremans, Li Su
Abstract
Most of the current supervised automatic music transcription (AMT) models lack the ability to generalize. This means that they have trouble transcribing real-world music recordings from diverse musical genres that are not presented in the labelled training data. In this paper, we propose a semi-supervised framework, ReconVAT, which solves this issue by leveraging the huge amount of available unlabelled music recordings. The proposed ReconVAT uses reconstruction loss and virtual adversarial training. When combined with existing U-net models for AMT, ReconVAT achieves competitive results on common benchmark datasets such as MAPS and MusicNet. For example, in the few-shot setting for the string part version of MusicNet, ReconVAT achieves F1-scores of 61.0% and 41.6% for the note-wise and note-with-offset-wise metrics respectively, which translates into an improvement of 22.2% and 62.5% compared to the supervised baseline model. Our proposed framework also demonstrates the potential of continual learning on new data, which could be useful in real-world applications whereby new data is constantly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1d2dc37-cfcd-4108-813c-fe90dcdaa942Cited by top-tier papers3
- MT3: Multi-Task Multitrack Music TranscriptionJosh Gardner, Ian Simon, Ethan Manilow, Curtis Hawthorne et al.ICLR 2022 · 134 citations
- Unaligned Supervision for Automatic Music Transcription in The WildBen Maman, Amit H. BermanoICML 2022 · 46 citations
- Text2Data: Low-Resource Data Generation with Textual ControlShiyu Wang, Yihao Feng, Tian Lan, Ning Yu et al.AAAI 2025
Builds on2
Related papers
- MCSSME: Multi-Task Contrastive Learning for Semi-supervised Singing Melody Extraction from Polyphonic MusicShuai YuAAAI 2024 · 9 citations
- Robust Singing Voice Transcription Serves SynthesisRuiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong et al.ACL 2024 · 6 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured TokenizationLongshen Ou, Jingwei Zhao, Ziyu Wang, Gus Xia et al.NeurIPS 2025 · 5 citations
- MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingYizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma et al.ICLR 2024 · 277 citations
