Lune

ACM MM2025Top-tier venue

Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation

Weitao You, Heda Zuo, Junxian Wu, Dengming Zhang, Zhibin Zhou, Lingyun Sun

2025Year

Abstract

Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines