Fine-grained Audible Video Description
Xuyang Shen, Dong Li, Jinxing Zhou, Zhen Qin, Bowen He, Xiaodong Han, Aixuan Li, Yuchao Dai, Lingpeng Kong, Meng Wang, Yu Qiao, Yiran Zhong
Abstract
We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of each object, the actions of moving objects, and the sounds in videos. Existing visual-language modeling tasks often concentrate on visual cues in videos while undervaluing the language and audio modalities. On the other hand, FAVD requires not only audio-visual-language modeling skills but also paragraph-level language generation abilities. We construct the first fine-grained audible video description benchmark (FAVDBench) to facilitate this research. For each video clip, we first provide a one-sentence summary of the video, i.e., the caption, followed by 4-6 sentences describing the visual details and 1-2 audio-related descriptions at the end. The descriptions are provided in both English and Chinese. We create two new metrics for this task: an EntityScore to gauge the completeness of entities in the visual descriptions, and an Au-dioScore to assess the audio descriptions. As a preliminary approach to this task, we propose an audio-visuallanguage transformer that extends existing video captioning model with an additional audio branch. We combine the masked language modeling and auto-regressive language modeling losses to optimize our model so that it can produce paragraph-level descriptions. We illustrate the efficiency of our model in audio-visual-language modeling by evaluating it against the proposed benchmark using both conventional captioning metrics and our proposed metrics. We further put our benchmark to the test in video generation models, demonstrating that employing finegrained video descriptions can create more intricate videos than using captions. Code and dataset are available at https://github.com/OpenNLPLab/FAVDBench . Our online benchmark is available at www.avlbench.opennlplab.cn.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03d463a2-dc81-49ea-95e8-248d28af006bCited by top-tier papers15
- JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior SynchronizationKai Liu, Wei Li, Lai Chen, Shengqiong Wu et al.ICLR 2026 · 89 citations
- Object-Aware Adaptive-Positivity Learning for Audio-Visual Question AnsweringZhangbin Li, Dan Guo, Jinxing Zhou, Jing Zhang et al.AAAI 2024 · 30 citations
- Patch-level Sounding Object Tracking for Audio-Visual Question AnsweringZhangbin Li, Jinxing Zhou, Jing Zhang, Shengeng Tang et al.AAAI 2025 · 20 citations
- Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video ParsingPengcheng Zhao, Jinxing Zhou, Yang Zhao, Dan Guo et al.AAAI 2025 · 19 citations
- TAVGBench: Benchmarking Text to Audible-Video GenerationYuxin Mao, Xuyang Shen, Jing Zhang, Zhen Qin et al.ACM MM 2024 · 12 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video GenerationZiwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang et al.ICML 2026 · 8 citations
- FineVAU: A Novel Human-Aligned Benchmark for Fine-Grained Video Anomaly UnderstandingJoão Alexandre Cardeira Pereira, Vasco Lopes, João Neves, David SemedoAAAI 2026
- DeVAn: Dense Video Annotation for Video-Language ModelsTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan et al.ACL 2024 · 1 citation
- MAVT-FG: Multimodal Audio-Visual Transformer for Weakly-supervised Fine-Grained RecognitionXiaoyu Zhou, Xiaotong Song, Hao Wu, Jingran Zhang et al.ACM MM 2022 · 4 citations
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New BenchmarkWenhao Chai, Enxin Song, Yilun Du, Chenlin Meng et al.ICLR 2025
