Hierarchical Modular Network for Video Captioning
Hanhua Ye, Guorong Li, Yuankai Qi, Shuhui Wang, Qingming Huang, Ming-Hsuan Yang
摘要
Video captioning aims to generate natural language descriptions according to the content, where representation learning plays a crucial role. Existing methods are mainly developed within the supervised learning framework via word-by-word comparison of the generated caption against the ground-truth text without fully exploiting linguistic semantics. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics from three levels before generating captions. In particular, the hierarchy is composed of: (I) Entity level, which highlights objects that are most likely to be mentioned in captions. (II) Predicate level, which learns the actions conditioned on highlighted objects and is supervised by the predicate in captions. (III) Sentence level, which learns the global semantic representation and is supervised by the whole caption. Each level is implemented by one module. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on the two widely-used benchmarks: MSVD 104.0% and MSR-VTT 51.5% in CIDEr score. Code will be made available at https://github.com/MarcusNerva/HMN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Refined Semantic Enhancement towards Frequency Diffusion for Video CaptioningXian Zhong, Zipeng Li, Shuqin Chen, Kui Jiang 等AAAI 2023 · 被引用 70 次
- Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable SensorsWenqiang Chen, Jiaxuan Cheng, Leyao Wang, Wei Zhao 等UbiComp 2025 · 被引用 24 次
- Exploring Group Video Captioning with Efficient Relational ApproximationWang Lin, Tao Jin, Ye Wang, Wenwen Pan 等ICCV 2023 · 被引用 17 次
- RTQ: Rethinking Video-language Understanding Based on Image-text ModelXiao Wang, Yaoyu Li, Tian Gan, Zheng Zhang 等ACM MM 2023 · 被引用 14 次
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
相关 Paper
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 被引用 13 次
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu 等ACM MM 2021 · 被引用 26 次
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Comprehending and Ordering Semantics for Image CaptioningYehao Li, Yingwei Pan, Ting Yao, Tao MeiCVPR 2022 · 被引用 124 次
- Comprehensive Visual Grounding for Video DescriptionWenhui Jiang, Yibo Cheng, Linxin Liu, Yuming Fang 等AAAI 2024 · 被引用 5 次
