Audeo: Audio Generation for a Silent Performance Video
Kun Su, Xiulong Liu, Eli Shlizerman
Abstract
We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named *Audeo*' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. *Audeo* converts video to audio smoothly and clearly with only a few setup constraints. We evaluate *Audeo* on in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1115c9d6-0eee-4706-84c6-ee44fa9e4cceCited by top-tier papers22
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 92 citations
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang et al.ACM MM 2021 · 87 citations
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang et al.NeurIPS 2023 · 60 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
Related papers
- PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano PerformanceQijun Gan, Song Wang, Shengtao Wu, Jianke ZhuICLR 2025 · 1 citation
- Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key MotionJingjing Tang, Shinichi Furuya, Hayato Nishioka, Momoko Shioki et al.CHI 2026 · 1 citation
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin et al.AAAI 2024 · 29 citations
- ReTouche: Embodied Representations for Self-Guided Piano LearningPaul-Peter Arslan, Hayoun Noh, Mariana Aki Tamashiro, Louis Badr et al.CHI 2026 · 1 citation
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia et al.ICML 2026 · 3 citations
