Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
Teng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su, Zhengguang Zhou, Youliang Zhang, Yuan Zhou, Qinglin Lu, Ran Yi
摘要
The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift, where concurrently evolving noisy latents impede stable learning of alignment; (2) inefficient global attention mechanisms that fail to capture fine-grained temporal cues; and (3) the intra-modal bias of conventional Classifier-Free Guidance (CFG), which enhances conditionality but not cross-modal synchronization. To overcome these challenges, we introduce Harmony, a novel framework that mechanistically enforces audio-visual synchronization. We first propose a Cross-Task Synergy training paradigm to mitigate drift by leveraging strong supervisory signals from audio-driven video and video-driven audio generation tasks. Then, we design a Global-Local Decoupled Interaction Module for efficient and precise temporal-style alignment. Finally, we present a novel Synchronization-Enhanced CFG (SyncCFG) that explicitly isolates and amplifies the alignment signal during inference. Extensive experiments demonstrate that Harmony establishes a new state-of-the-art, significantly outperforming existing methods in both generation fidelity and, critically, in achieving fine-grained audio-visual synchronization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video GenerationKai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu 等ICLR 2026 · 被引用 24 次
- T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video GenerationZhe Cao, Tao Wang, Jiaming Wang, Yanghai Wang 等ICML 2026 · 被引用 13 次
- MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video GenerationYanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang 等ACL 2026 · 被引用 7 次
- PoseAnything: General Pose-guided Video Generation with Part-aware Temporal CoherenceRuiyan Wang, Teng Hu, Kaihui Huang, Zihan Su 等CVPR 2026
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng 等CVPR 2026 · 被引用 40 次
- Human-Centric Video Generation via Collaborative Multi-Modal ConditioningLiyang Chen, Tianxiang Ma, Jiawei Liu, Bingchuan Li 等AAAI 2026 · 被引用 1 次
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 等ICCV 2025 · 被引用 3 次
- Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio GenerationKang Zhang, Trung X. Pham, Suyeon Lee, Axi Niu 等NeurIPS 2025 · 被引用 1 次
- Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent AlignersYazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 等CVPR 2024 · 被引用 25 次
