Audeo: Audio Generation for a Silent Performance Video
Kun Su, Xiulong Liu, Eli Shlizerman
摘要
We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named *Audeo*' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. *Audeo* converts video to audio smoothly and clearly with only a few setup constraints. We evaluate *Audeo* on in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 等ICML 2023 · 被引用 469 次
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang 等ACM MM 2021 · 被引用 87 次
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang 等NeurIPS 2023 · 被引用 60 次
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao 等ICCV 2023 · 被引用 51 次
相关 Paper
- PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano PerformanceQijun Gan, Song Wang, Shengtao Wu, Jianke ZhuICLR 2025 · 被引用 1 次
- Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key MotionJingjing Tang, Shinichi Furuya, Hayato Nishioka, Momoko Shioki 等CHI 2026 · 被引用 1 次
- V2Meow: Meowing to the Visual Beat via Video-to-Music GenerationKun Su, Judith Yue Li, Qingqing Huang, Dima Kuzmin 等AAAI 2024 · 被引用 29 次
- ReTouche: Embodied Representations for Self-Guided Piano LearningPaul-Peter Arslan, Hayoun Noh, Mariana Aki Tamashiro, Louis Badr 等CHI 2026 · 被引用 1 次
- Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-TrainingHong-Jie You, Jie-Jing Shao, Xiao-Wen Yang, Lin-Han Jia 等ICML 2026 · 被引用 3 次
