EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation
Bingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang, Shu Zhang, Longyin Wen, Yulei Niu
Abstract
Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-to-audio (VT2A), the current formulation faces three key limitations: (1) an imbalance between visual and textual conditioning that leads to visual dominance; (2) the absence of a concrete definition for finegrained controllable generation; (3) weak instruction understanding and following, as existing datasets rely on brief categorical tags. To address these limitations, we introduce EchoFoley (Event-Centric Hierarchical cOntrol), a new task designed for video-grounded sound generation with both event-level local control and hierarchical semantic control. Our symbolic representation for sounding events specifies when, what, and how each sound is produced within a video or instruction, enabling finegrained controls like sound generation, insertion, and editing. To support this task, we construct EchoFoley-6k, a large-scale, expert-curated benchmark containing over 6,000 video-instruction-annotation triplets and 42,000 fine-grained sounding event annotations. Building upon this foundation, we propose EchoVidia, a sounding-eventcentric agentic generation framework with slow-fast thinking strategy. Experiments show that EchoVidia surpasses recent VT2A models by 40.7% in controllability and 12.5% in perceptual quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e776f82f-4c0c-4dd8-af90-a08558f27b84Builds on19
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
- Perception Encoder: The best visual embeddings are not at the output of the networkDaniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho et al.NeurIPS 2025 · 359 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion ModelsSimian Luo, Chuanhao Yan, Chenxu Hu, Hang ZhaoNeurIPS 2023 · 192 citations
Related papers
- AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic TransferPengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen et al.ICLR 2026 · 3 citations
- CAFA: A Controllable Automatic Foley ArtistRoi Benita, Michael Finkelson, Tavi Halperin, Gleb Sterkin et al.ICCV 2025 · 1 citation
- FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured ScriptsYou Li, Dewei Zhou, Fan Ma, Fu Li et al.CVPR 2026 · 2 citations
- ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and EditingHuadai Liu, Kaicheng Luo, Jialei Wang, Wen Wang et al.NeurIPS 2025 · 5 citations
- AudioX: A Unified Framework for Anything-to-Audio GenerationZeyue Tian, Zhaoyang Liu, Yizhu Jin, Ruibin Yuan et al.ICLR 2026 · 38 citations
