Spatiotemporal Fine-grained Video Description for Short Videos
Te Yang, Jian Jia, Bo Wang, Yanhua Cheng, Yan Li, Dongze Hao, Xipeng Cao, Quan Chen, Han Li, Peng Jiang, Xiangyu Zhu, Zhen Lei
Abstract
In the mobile internet era, short videos are inundating people's lives. However, research on visual language models specifically designed for short videos has not yet received sufficient attention. Short videos are not just videos of limited duration. The prominent visual details and high information density of short videos differentiate them to long videos. In this paper, we propose the SpatioTemporal Fine-grained Description (STFVD) emphasizing on the uniqueness of short videos, which entails capturing the intricate details of the main subject and fine-grained movements. To this end, we create a comprehensive Short Video Advertisements Description (SVAD) dataset, comprising 34,930 clips from 5,046 videos. The dataset covers a range of topics, including 191 sub-industries, 649 popular products, and 470 trending games. Various efforts have been made in the data annotation process to ensure the inclusion of fine-grained spatiotemporal information, resulting in 34,930 high-quality annotations. Compared to existing datasets, samples in SVAD exhibit a superior text information density, suggesting that SVAD is more appropriate for the analysis of short videos. Based on the SVAD dataset, we develop a visual language model (SVAD-VLM) to generate spatiotemporal fine-grained description for short videos. We use a prompt-guided keyword generation task to efficiently learn key visual information. Moreover, we also utilize dual visual alignment to exploit the advantage of mixed-datasets training. Experiments on SVAD dataset demonstrate the challenge of STFVD and the competitive performance of proposed method compared to previous ones.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4dc7297c-579e-4c87-965f-06fe8852cf92Related papers
- Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal PromptsPeng Wu, Xuerong Zhou, Guansong Pang, Zhiwei Yang et al.ACM MM 2024 · 50 citations
- A Large-Scale Dataset for Short-Video Topic Peak Prediction and a Large Heterogeneous Graph ModelShangheng Chen, Shengsheng Qian, Quan Fang, Jun Hu et al.ACM MM 2025
- DeVAn: Dense Video Annotation for Video-Language ModelsTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan et al.ACL 2024 · 1 citation
- 3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short VideosVikram Gupta, Trisha Mittal, Puneet Mathur, Vaibhav Mishra et al.CVPR 2022 · 14 citations
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingYangliu Hu, Zikai Song, Na Feng, Yawei Luo et al.CVPR 2025
