FilmComposer: LLM-Driven Music Production for Silent Film Clips
Zhifeng Xie, Qile He, Youjia Zhu, Qiwei He, Mengtian Li
Abstract
In this work, we implement music production for silent film clips using LLM-driven method. Given the strong professional demands of film music production, we propose the FilmComposer, simulating the actual workflows of professional musicians. FilmComposer is the first to combine large generative models with a multi-agent approach, leveraging the advantages of both waveform music and symbolic music generation. Additionally, FilmComposer is the first to focus on the three core elements of music production for film-audio quality, musicality, and musical development-and introduces various controls, such as rhythm, semantics, and visuals, to enhance these key aspects. Specifically, FilmComposer consists of the visual processing module, rhythm-controllable MusicGen, and multi-agent assessment, arrangement and mix. In addition, our framework can seamlessly integrate into the actual music production pipeline and allows user intervention in every step, providing strong interactivity and a high degree of creative freedom. Furthermore, we propose MusicPro-7k which includes 7,418 film clips, music, description, rhythm spots and main melody, considering the lack of a professional and high-quality film music dataset. Finally, both the standard metrics and the new specialized metrics we propose demonstrate that the music generated by our model achieves state-of-the-art performance in terms of quality, consistency with video, diversity, musicality, and musical development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 410ce94d-a717-4ca4-b084-ae87d86e0676Cited by top-tier papers3
- Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music GenerationXinyi Tong, Yiran Zhu, Jishang Chen, Chunru Zhan et al.AAAI 2026 · 4 citations
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio GenerationYan Rong, Jinting Wang, Guangzhi Lei, Shan Yang et al.ACM MM 2025 · 1 citation
- VidTune: Creating Video Soundtracks with Generative Music and Video-Based ThumbnailsMina Huh, C. Ailie Fraser, Dingzeyu Li, Mira Dontcheva et al.CHI 2026 · 1 citation
Builds on13
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- Video Background Music Generation with Controllable Music TransformerShangzhe Di, Zeren Jiang, Si Liu, Zhaokai Wang et al.ACM MM 2021 · 87 citations
- Video Background Music Generation: Dataset, Method and EvaluationLe Zhuo, Zhaokai Wang, Baisen Wang, Yue Liao et al.ICCV 2023 · 51 citations
Related papers
- WildScore: Benchmarking MLLMs in-the-Wild Symbolic Music ReasoningGagan Mundada, Yash Vishe, Amit Namburi, Xin Xu et al.EMNLP 2025 · 1 citation
- AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality AssessmentYuqin Cao, Xiongkuo Min, Yixuan Gao, Wei Sun et al.ICML 2025
- FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film ClipsMengtian Li, Kunyan Dai, Yi Ding, Ruobing Ni et al.CVPR 2026 · 1 citation
- AutoAD III: The Prequel - Back to the PixelsTengda Han, Max Bain, Arsha Nagrani, Gül Varol et al.CVPR 2024
- SongComposer: A Large Language Model for Lyric and Melody Generation in Song CompositionShuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang et al.ACL 2025
