Unaligned Supervision for Automatic Music Transcription in The Wild
Ben Maman, Amit H. Bermano
摘要
Multi-instrument Automatic Music Transcription (AMT), or the decoding of a musical recording into semantic musical content, is one of the holy grails of Music Information Retrieval. Current AMT approaches are restricted to piano and (some) guitar recordings, due to difficult data collection. In order to overcome data collection barriers, previous AMT approaches attempt to employ musical scores in the form of a digitized version of the same song or piece. The scores are typically aligned using audio features and strenuous human intervention to generate training labels. We introduce Note EM , a method for simultaneously training a transcriber and aligning the scores to their corresponding performances, in a fully-automated process. Using this unaligned supervision scheme, complemented by pseudolabels and pitch-shift augmentation, our method enables training on in-the-wild recordings with unprecedented accuracy and instrumental variety. Using only synthetic data and unaligned supervision, we report SOTA note-level accuracy of the MAPS dataset, and large favorable margins on cross-dataset evaluations. We also demonstrate robustness and ease of use; we report comparable results when training on a small, easily obtainable, self-collected dataset, and we propose alternative labeling to the MusicNet dataset, which we show to be more accurate. Our project page is available at https://benadar293.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- MUSCAT: A Multimodal mUSic Collection for Automatic Transcription of Real Recordings and Image ScoresAlejandro Galán-Cuenca, Jose J. Valero-Mas, Juan C. Martinez-Sevilla, Antonio Hidalgo-Centeno 等ACM MM 2024 · 被引用 2 次
- Bridging Piano Transcription and Rendering via Disentangled Score Content and StyleWei Zeng, Junchuan Zhao, Ye WangICLR 2026
- Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured TokenizationLongshen Ou, Jingwei Zhao, Ziyu Wang, Gus Xia 等NeurIPS 2025 · 被引用 5 次
- Detecting Music Performance Errors with TransformersBenjamin Shiue-Hal Chou, Purvish Jajal, Nicholas John Eliopoulos, Tim Nadolsky 等AAAI 2025 · 被引用 3 次
- Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic MusicHongju Su, Ke Li, Lan Yang, Honggang Zhang 等ACL 2026
