Spatiotemporal Fine-grained Video Description for Short Videos
Te Yang, Jian Jia, Bo Wang, Yanhua Cheng, Yan Li, Dongze Hao, Xipeng Cao, Quan Chen, Han Li, Peng Jiang, Xiangyu Zhu, Zhen Lei
摘要
In the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang 等ACM MM 2024 · 被引用 50 次
- A Large-Scale Dataset for Short-Video Topic Peak Prediction and a Large Heterogeneous Graph ModelShangheng Chen, Shengsheng Qian, Quan Fang, Jun Hu 等ACM MM 2025
- DeVAn: Dense Video Annotation for Video-Language ModelsTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan 等ACL 2024 · 被引用 1 次
- 3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short VideosVikram Gupta, Trisha Mittal, Puneet Mathur, Vaibhav Mishra 等CVPR 2022 · 被引用 14 次
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingYangliu Hu, Zikai Song, Na Feng, Yawei Luo 等CVPR 2025
