Poet: Product-oriented Video Captioner for E-commerce
Shengyu Zhang, Ziqi Tan, Jin Yu, Zhou Zhao, Kun Kuang, Jie Liu, Jingren Zhou, Hongxia Yang, Fei Wu
Abstract
In e-commerce, a growing number of user-generated videos are used for product promotion. How to generate video descriptions that narrate the user-preferred product characteristics depicted in the video is vital for successful promoting. Traditional video captioning methods, which focus on routinely describing what exists and happens in a video, are not amenable for product-oriented video captioning. To address this problem, we propose a product-oriented video captioner framework, abbreviated as Poet. Poet firstly represents the videos as product-oriented spatial-temporal graphs. Then, based on the aspects of the video-associated product, we perform knowledge-enhanced spatial-temporal inference on those graphs for capturing the dynamic change of fine-grained product-part characteristics. The knowledge leveraging module in Poet differs from the traditional design by performing knowledge filtering and dynamic memory modeling. We show that Poet achieves consistent performance improvement over previous methods concerning generation quality, product aspects capturing, and lexical diversity. Experiments are performed on two product-oriented video captioning datasets, buyer-generated fashion video dataset (BFVD) and fan-generated fashion video dataset (FFVD), collected from Mobile Taobao. We will release the desensitized datasets to promote further investigations on both video captioning and general video analysis problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 332ea764-dbeb-4843-a331-150139fc9fb2Cited by top-tier papers11
- Re4: Learning to Re-contrast, Re-attend, Re-construct for Multi-interest RecommendationShengyu Zhang, Lingxiao Yang, Dong Yao, Yujie Lu et al.WWW 2022 · 67 citations
- SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationRongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren et al.ACM MM 2022 · 46 citations
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
- Future-Aware Diverse Trends Framework for RecommendationYujie Lu, Shengyu Zhang, Yingxuan Huang, Luyao Wang et al.WWW 2021 · 38 citations
- Adaptive Hierarchical Graph Reasoning with Semantic Coherence for Video-and-Language InferenceJuncheng Li, Siliang Tang, Linchao Zhu, Haochen Shi et al.ICCV 2021 · 28 citations
Builds on2
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
Related papers
- Edit As You Wish: Video Caption Editing with Multi-grained User ControlLinli Yao, Yuanmeng Zhang, Ziheng Wang, Xinglin Hou et al.ACM MM 2024 · 4 citations
- Syntax-Aware Action Targeting for Video CaptioningQi Zheng, Chaoyue Wang, Dacheng TaoCVPR 2020
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu et al.ACM MM 2021 · 26 citations
- Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and BenchmarksZixuan Xiong, Guangwei Xu, Wenkai Zhang, Yuan Miao et al.ICLR 2025
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding et al.ACM MM 2021 · 14 citations
