Skipping the Frame-Level: Event-Based Piano Transcription With Neural Semi-CRFs
Yujia Yan, Frank Cwitkowitz, Zhiyao Duan
摘要
Piano transcription systems are typically optimized to estimate pitch activity at each frame of audio. They are often followed by carefully designed heuristics and post-processing algorithms to estimate note events from the frame-level predictions. Recent methods have also framed piano transcription as a multi-task learning problem, where the activation of different stages of a note event are estimated independently. These practices are not well aligned with the desired outcome of the task, which is the specification of note intervals as holistic events, rather than the aggregation of disjoint observations. In this work, we propose a novel formulation of piano transcription, which is optimized to directly predict note events. Our method is based on Semi-Markov Conditional Random Fields (semi-CRF), which produce scores for intervals rather than individual frames. When formulating piano transcription in this way, we eliminate the need to rely on disjoint frame-level estimates for different stages of a note event. We conduct experiments on the MAESTRO dataset and demonstrate that the proposed model surpasses the current state-of-the-art for piano transcription. Our results suggest that the semi-CRF output layer, while still quadratic in complexity, is a simple, fast and well-performing solution for event-based prediction, and may lead to similar success in other areas which currently rely on frame-level estimates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- MAJL: A Model-Agnostic Joint Learning Framework for Music Source Separation and Pitch EstimationHaojie Wei, Jun Yuan, Rui Zhang, Quanyu Dai 等ACM MM 2024 · 被引用 3 次
- Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music ModelingLouis Bradshaw, Simon ColtonICLR 2025
它引用的顶会 Paper1
相关 Paper
- Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key MotionJingjing Tang, Shinichi Furuya, Hayato Nishioka, Momoko Shioki 等CHI 2026 · 被引用 1 次
- Automatic Piano Fingering from Partially Annotated Scores using Autoregressive Neural NetworksPedro Ramoneda, Dasaem Jeong, Eita Nakamura, Xavier Serra 等ACM MM 2022 · 被引用 7 次
- Robust Singing Voice Transcription Serves SynthesisRuiqi Li, Yu Zhang, Yongqi Wang, Zhiqing Hong 等ACL 2024 · 被引用 6 次
- Unaligned Supervision for Automatic Music Transcription in The WildBen Maman, Amit H. BermanoICML 2022 · 被引用 46 次
- Bridging Piano Transcription and Rendering via Disentangled Score Content and StyleWei Zeng, Junchuan Zhao, Ye WangICLR 2026
