Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs
Yujia Yan, Frank Cwitkowitz, Zhiyao Duan
Abstract
Piano transcription systems are typically optimized to estimate pitch activity at each frame of audio. They are often followed by carefully designed heuristics and post-processing algorithms to estimate note events from the frame-level predictions. Recent methods have also framed piano transcription as a multi-task learning problem, where the activation of different stages of a note event are estimated independently. These practices are not well aligned with the desired outcome of the task, which is the specification of note intervals as holistic events, rather than the aggregation of disjoint observations. In this work, we propose a novel formulation of piano transcription, which is optimized to directly predict note events. Our method is based on Semi-Markov Conditional Random Fields (semi-CRF), which produce scores for intervals rather than individual frames. When formulating piano transcription in this way, we eliminate the need to rely on disjoint frame-level estimates for different stages of a note event. We conduct experiments on the MAESTRO dataset and demonstrate that the proposed model surpasses the current state-of-the-art for piano transcription. Our results suggest that the semi-CRF output layer, while still quadratic in complexity, is a simple, fast and well-performing solution for event-based prediction, and may lead to similar success in other areas which currently rely on frame-level estimates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch EstimationHaojie Wei, Jun Yuan, Rui Zhang, Quanyu Dai et al.ACM MM 2024 · 3 citations
- Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music ModelingLouis Bradshaw, Simon ColtonICLR 2025
Builds on1
Related papers
- Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key MotionJingjing Tang, Shinichi Furuya, Hayato Nishioka, Momoko Shioki et al.CHI 2026 · 1 citation
- Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural NetworksPedro Ramoneda, Dasaem Jeong, Eita Nakamura, Xavier Serra et al.ACM MM 2022 · 7 citations
- Robust Singing Voice Transcription Serves SynthesisRuiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong et al.ACL 2024 · 6 citations
- Unaligned Supervision for Automatic Music Transcription in The WildBen Maman, Amit H. BermanoICML 2022 · 46 citations
- Bridging Piano Transcription and Rendering via Disentangled Score Content and StyleWei Zeng, Junchuan Zhao, Ye WangICLR 2026
