Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
Hong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia, Lan-Zhe Guo, Yu-Feng Li
Abstract
Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance from a symbolic score, rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language. To address this gap, we introduce Pianist Transformer, with three key contributions: 1) introducing large-scale selfsupervised learning into expressive piano performance rendering through a unified Musical Instrument Digital Interface (MIDI) representation, enabling pre-training on 10B tokens of unlabeled MIDI data; 2) an efficient asymmetric Transformer with note-level compression, substantially improving training efficiency, memory usage, and inference speed for long-context music modeling; 3) a state-of-the-art rendering model with an editable workflow, achieving strong objective and subjective results and enabling integration into real-world music production workflows. Overall, Pianist Transformer outlines a scalable path toward human-like performance synthesis in the music domain. Code, audio samples, and model checkpoints are available on our project page: https://yhj137.github. io/pianist-transformer-demo/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- On the Learnability of Test-Time Adaptation: A Recovery Complexity PerspectiveZhi Zhou, Ming Yang, Shi-Yu Tian, Kun-Yang Yu et al.ICML 2026 · 2 citations
- AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music EditingChih-Heng Chang, Keng-Seng Ho, Chih-Yu Tsai, Kuan-Lin Chen et al.KDD 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 242 citations
Related papers
- Bridging Piano Transcription and Rendering via Disentangled Score Content and StyleWei Zeng, Junchuan Zhao, Ye WangICLR 2026
- SyMuPe: Affective and Controllable Symbolic Music PerformanceIlya Borovik, Dmitrii Gavrilev, Vladimir ViroACM MM 2025
- Aria-MIDI: A Dataset of Piano MIDI Files for Symbolic Music ModelingLouis Bradshaw, Simon ColtonICLR 2025
- N-gram Unsupervised Compoundation and Feature Injection for Better Symbolic Music UnderstandingJinhao Tian, Zuchao Li, Jiajia Li, Ping WangAAAI 2024
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu et al.ICML 2020 · 102 citations
