AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
Xinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua, Linli Yao, Yang Shi, Bozhou Li, Qiang Liu, Yuanxing Zhang, Pengfei Wan, Liang Wang
摘要
Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a powerful audiovisual video captioner driven by the temporal orchestration between audio and visual modalities. We propose a two-stage post-training pipeline: (1) AVoCaDO SFT, which fine-tunes the model on a newly curated dataset of 107K high-quality, temporally-aligned audiovisual captions; and (2) AVoCaDO GRPO, which leverages tailored reward functions to further enhance temporal coherence and dialogue accuracy while regularizing caption length and reducing collapse. Experimental results demonstrate that AVoCaDO significantly outperforms existing open-source models across four audiovisual video captioning benchmarks, and also achieves competitive performance on the VDC benchmark under visual-only settings. The model will be made publicly available to facilitate future research in audiovisual video understanding and generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language ModelsYue Ding, Yiyan Ji, Jungang Li, Xuyang Liu 等ICML 2026 · 被引用 22 次
- FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMsQian Chen, Jinlan Fu, Changsong Li, Min zhang 等ICML 2026 · 被引用 5 次
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information LossBozhou Li, Xinda Xue, Sihan Yang, Yang Shi 等ICLR 2026 · 被引用 5 次
- TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual CaptionsLinli Yao, Yuancheng Wei, Yaojie Zhang, Lei Li 等ICML 2026
它引用的顶会 Paper17
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du 等NeurIPS 2025 · 被引用 143 次
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersWenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu 等ICLR 2023 · 被引用 116 次
相关 Paper
- Learning to See through Sound: From VggCaps to Multi2Cap for Richer Automated Audio CaptioningSangyeon Cho, Mingi Kim, Jinkwon Hwang, Jaehoon Go 等EMNLP 2025
- An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matchingHugo Malard, Michel Olvera, Stéphane Lathuilière, Slim EssidNeurIPS 2024 · 被引用 3 次
- Fine-grained Audible Video DescriptionXuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin 等CVPR 2023
- Aligned Better, Listen Better for Audio-Visual Large Language ModelsYuxin Guo, Shuailei Ma, Shijie Ma, Xiaoyi Bao 等ICLR 2025
- Towards Fine-grained Audio Captioning with Multimodal Contextual FusionShunian Chen, Xinyuan Xie, Zheshu Chen, Owen Lee 等ACL 2026
