Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation
Jiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang, Dadong Wang, Tongliang Liu
Abstract
Text-to-Audio-Video (T2AV) generation aims to produce temporally and semantically aligned visual and auditory content from natural language descriptions. While recent progress in text-to-audio and text-to-video models has improved generation quality within each modality, jointly modeling them remains challenging due to incomplete and asymmetric correspondence: audio often reflects only a subset of the visual scene, and vice versa. Naively enforcing full alignment introduces semantic noise and temporal mismatches. To address this, we propose a novel framework that performs selective cross-modal alignment through a learnable masking mechanism, enabling the model to isolate and align only the shared latent components relevant to both modalities. This mechanism is integrated into an adaptation module that interfaces with pretrained encoders and decoders from latent video and audio diffusion models, preserving their generative capacity with reduced training overhead. Theoretically, we show that our masked objective provably recovers the minimal set of shared latent variables across modalities. Empirically, our method achieves state-of-the-art performance on standard T2AV benchmarks, demonstrating significant improvements in audiovisual synchronization and semantic consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov et al.ICCV 2025 · 3 citations
- TiVA: Time-Aligned Video-to-Audio GenerationXihua Wang, Yuyue Wang, Yihan Wu, Ruihua Song et al.ACM MM 2024 · 6 citations
- TAVGBench: Benchmarking Text to Audible-Video GenerationYuxin Mao, Xuyang Shen, Jing Zhang, Zhen Qin et al.ACM MM 2024 · 12 citations
- Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersYazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang et al.CVPR 2024 · 25 citations
- MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio SynthesisHo Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya et al.CVPR 2025
