Hierarchical Context-aware Network for Dense Video Event Captioning
Lei Ji, Xianglin Guo, Haoyang Huang, Xilin Chen
摘要
Dense video event captioning aims to generate a sequence of descriptive captions for each event in a long untrimmed video. Video-level context provides important information and facilities the model to generate consistent and less redundant captions between events. In this paper, we introduce a novel Hierarchical Context-aware Network for dense video event captioning (HCN) to capture context from various aspects. In detail, the model leverages local and global context with different mechanisms to jointly learn to generate coherent captions. The local context module performs full interaction between neighbor frames and the global context module selectively attends to previous or future events. According to our extensive experiment on both Youcook2 and Activitynet Captioning datasets, the videolevel HCN model outperforms the event-level context-agnostic model by a large margin. The code is available at https://github.com/ KirkGuo/HCN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation LearningSimon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxNeurIPS 2020 · 被引用 186 次
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
- Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event CaptioningTanzila Rahman, Bicheng Xu, Leonid SigalICCV 2019 · 被引用 89 次
相关 Paper
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang 等AAAI 2025 · 被引用 5 次
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 等CVPR 2024
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li 等AAAI 2026 · 被引用 1 次
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- Event-Equalized Dense Video CaptioningKangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li 等CVPR 2025
