Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
Junwon Lee, Juhan Nam, Jiyoung Lee
摘要
This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio tracks are handled individually for each sound source for precise editing, mixing, and creative control. We propose SELVA, a novel text-conditioned V2A model that treats the text prompt as an explicit selector to distinctly extract prompt-relevant sound-source visual features from the video encoder. To suppress text-irrelevant activations with efficient video encoder finetuning, the proposed supplementary tokens promote cross-attention to yield robust semantic and temporal grounding. SELVA further employs an autonomous video-mixing scheme in a self-supervised manner to overcome the lack of mono audio track supervision. We evaluate SELVA on VGG-MONOAUDIO, a curated benchmark of clean single-source videos for such a task. Extensive experiments and ablations consistently verify its effectiveness across audio quality, semantic alignment, and temporal synchronization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- SelM: Selective Mechanism based Audio-Visual SegmentationJiaxu Li, Songsong Yu, Yifan Wang, Lijun Wang 等ACM MM 2024 · 被引用 5 次
- Tell What You Hear From What You See - Video to Audio Generation Through TextXiulong Liu, Kun Su, Eli ShlizermanNeurIPS 2024 · 被引用 46 次
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingXihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song 等ICCV 2025 · 被引用 6 次
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang 等NeurIPS 2025 · 被引用 6 次
- Gotta Hear Them All: Towards Sound Source Aware Audio GenerationWei Guo, Heng Wang, Jianbo Ma, Weidong CaiAAAI 2026 · 被引用 2 次
