AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov, Alper Canberk, Kwot Sin Lee, Vicente Ordonez, Sergey Tulyakov
Abstract
We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our framework is a Fusion Block that facilitates bidirectional information exchange between video and audio diffusion models through temporally-aligned self attention operations. Unlike prior work that uses dedicated models for A2V and V2A tasks and relies on pretrained feature extractors, AV-Link achieves both tasks in a single self-contained framework, directly leveraging features obtained by the complementary modality (i.e. video features to generate audio, or audio features to generate video). Extensive automatic and subjective evaluations demonstrate that our method achieves a substantial improvement in audio-video synchronization, outperforming more expensive baselines such as the MovieGen video-to-audio model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 688736e7-cf82-4ffd-9ea8-37b9c54947d2Cited by top-tier papers9
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyTeng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su et al.CVPR 2026 · 21 citations
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni et al.NeurIPS 2025 · 7 citations
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion TransformersMoayed Haji Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park et al.CVPR 2026 · 4 citations
- Benchmarking Single-Factor Physical Video-to-Audio GenerationTingle Li, Siddharth Gururani, Kevin Shih, Gantavya Bhatt et al.CVPR 2026 · 2 citations
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
Related papers
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang et al.NeurIPS 2025 · 6 citations
- Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersYazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang et al.CVPR 2024 · 25 citations
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingXihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song et al.ICCV 2025 · 6 citations
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song et al.ACM MM 2024 · 6 citations
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisHo Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya et al.CVPR 2025
