How Does it Sound?
Kun Su, Xiulong Liu, Eli Shlizerman
摘要
One of the primary purposes of video is to capture people and their unique activities. It is often the case that the experience of watching the video can be enhanced by adding a musical soundtrack that is in-sync with the rhythmic features of these activities. How would this soundtrack sound? Such a problem is challenging since little is known about capturing the rhythmic nature of free body movements. In this work, we explore this problem and propose a novel system, called 'RhythmicNet', which takes as an input a video with human movements and generates a soundtrack for it. RhythmicNet works directly with human movements, by extracting skeleton keypoints and implementing a sequence of models translating them to rhythmic sounds. RhythmicNet follows the natural process of music improvisation which includes the prescription of streams of the beat, the rhythm and the melody. In particular, RhythmicNet first infers the music beat and the style pattern from body keypoints per each frame to produce the rhythm. Next, it implements a transformerbased model to generate the hits of drum instruments and implements a U-net based model to generate the velocity and the offsets of the instruments. Additional types of instruments are added to the soundtrack by further conditioning on generated drum sounds. We evaluate RhythmicNet on large scale video datasets that include body movements with inherit sound association, such as dance, as well as 'in the wild' internet videos of various movements and actions. We show that the method can generate plausible music that aligns with different types of human movements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
- Connecting Multi-modal Contrastive RepresentationsZehan Wang, Yang Zhao, Xize Cheng, Haifeng Huang 等NeurIPS 2023 · 被引用 60 次
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao 等ICCV 2023 · 被引用 51 次
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 被引用 46 次
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 被引用 34 次
它引用的顶会 Paper8
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 被引用 271 次
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 被引用 265 次
- Encoding Musical Style with Transformer AutoencodersKristy Choi, Curtis Hawthorne, Ian Simon, Monica Dinculescu 等ICML 2020 · 被引用 102 次
- Audeo: Audio Generation for a Silent Performance VideoKun Su, Xiulong Liu, Eli ShlizermanNeurIPS 2020 · 被引用 78 次
- Scene-Aware Background Music SynthesisYujia Wang, Wei Liang, Wanwan Li, Dingzeyu Li 等ACM MM 2020 · 被引用 14 次
相关 Paper
- Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion ModelingXiaojie Li, Ronghui Li, Shukai Fang, Shuzhao Xie 等ICCV 2025 · 被引用 3 次
- Self-supervised Dance Video Synthesis Conditioned on MusicXuanchi Ren, Haoran Li, Zijian Huang, Qifeng ChenACM MM 2020 · 被引用 68 次
- Long-Term Rhythmic Video SoundtrackerJiashuo Yu, Yaohui Wang, Xinyuan Chen, Xiao Sun 等ICML 2023 · 被引用 24 次
- ChoreoGraph: Music-conditioned Automatic Dance Choreography over a Style and Tempo Consistent Dynamic GraphHo Yin Au, Jie Chen, Junkun Jiang, Yike GuoACM MM 2022 · 被引用 29 次
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationRonghui Li, Junfan Zhao, Yachao Zhang, Mingyang Su 等ICCV 2023 · 被引用 110 次
