Search-oriented Micro-video Captioning
Liqiang Nie, Leigang Qu, Dai Meng, Min Zhang, Qi Tian, Alberto Del Bimbo
Abstract
Pioneer efforts have been dedicated to the content-oriented video captioning that generates relevant sentences to describe the visual contents of a given video from the producer perspective. By contrast, this work targets at the search-oriented one that summarizes the given video via generating query-like sentences from the consumer angle. Beyond relevance, diversity is vital in characterizing consumers' seeking intention from different aspects. Towards this end, we devise a large-scale multimodal pre-training network regularized by five tasks to strengthen the downstream video representation, which is well-trained over our collected 11M micro-videos. Thereafter, we present a flow-based diverse captioning model to generate different captions from consumers' search demand. This model is optimized via a reconstruction loss and a KL divergence between the prior and the posterior. We justify our model over our constructed golden dataset comprising 690k pairs and experimental results demonstrate its superiority.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 74594809-a92c-4929-838d-cb3c1312e1a2Cited by top-tier papers16
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie et al.ACM MM 2023 · 91 citations
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li et al.ACM MM 2023 · 31 citations
- UniSA: Unified Generative Framework for Sentiment AnalysisZaijing Li, Ting-En Lin, Yuchuan Wu, Meng Liu et al.ACM MM 2023 · 22 citations
- Real20M: A Large-scale E-commerce Dataset for Cross-domain RetrievalYanzhe Chen, Huasong Zhong, Xiangteng He, Yuxin Peng et al.ACM MM 2023 · 15 citations
- RTQ: Rethinking Video-language Understanding Based on Image-text ModelXiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang et al.ACM MM 2023 · 14 citations
Related papers
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- End-to-end Generative Pretraining for Multimodal Video CaptioningPaul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia SchmidCVPR 2022 · 152 citations
- Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and BenchmarksZixuan Xiong, Guangwei Xu, Wenkai Zhang, Yuan Miao et al.ICLR 2025
- Open-Book Video Captioning With Retrieve-Copy-Generate NetworkZiqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan et al.CVPR 2021
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge et al.EMNLP 2022 · 5 citations
