Lune

ACM MM2025Top-tier venue

AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video Generation

Kai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos, Yapeng Tian

2025Year
2Citations
3Top-tier citations

Abstract

Recent Diffusion Transformers (DiTs) have shown impressive capabilities in generating single-modality content, including images, videos, and audio. However, the potential of DiTs to enable superb multimodal content creation remains underexplored. To bridge this gap, we introduce AV-DiT, a novel and efficient audio-visual diffusion transformer designed to generate high-quality, realistic videos with synchronized audio tracks. To minimize model complexity and computational costs, our AV-DiT utilizes a modality-shared DiT backbone pre-trained on image-only data, with only newly inserted adapters being trainable. This shared backbone facilitates the generation of both audio and video. Specifically, the video branch incorporates a trainable temporal attention layer into a pre-trained DiT block for capturing the temporal consistency for video generation. In addition, a small number of trainable parameters adapt the image-based DiT block to learn the acoustic characteristics for audio generation. An extra shared self-attention block reused from the DiT block, equipped with lightweight parameters, facilitates feature interaction between audio and visual modalities for alignment. Extensive experiments on the datasets demonstrate that our AV-DiT achieves state-of-the-art performance in joint audio-visual generation with significantly fewer tunable parameters. Furthermore, our results highlight that a single shared image generative backbone with modality-specific adaptations is sufficient for constructing a joint audio-video generator.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ae20d00c-600f-48da-a1b6-d83b27335dc0

Cited by top-tier papers3

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines