Learning Semantic Concepts and Temporal Alignment for Narrated Video Procedural Captioning
Botian Shi, Lei Ji, Zhendong Niu, Nan Duan, Ming Zhou, Xilin Chen
Abstract
Video captioning is a fundamental task for visual understanding. Previous works employ end-to-end networks to learn from the low-level vision feature and generate descriptive captions, which are hard to recognize fine-grained objects and lacks the understanding of crucial semantic concepts. According to DPC [19], these concepts generally present in the narrative transcripts of the instructional videos. The incorporation of transcript and video can improve the captioning performance. However, DPC directly concatenates the embedding of transcript with video features, which is incapable of fusing language and vision features effectively and leads to the temporal mis-alignment between transcript and video. This motivates us to 1) learn the semantic concepts explicitly and 2) design a temporal alignment mechanism to better align the video and transcript for the captioning task. In this paper, we start with an encoder-decoder backbone using transformer models. Firstly, we design a semantic concept prediction module as a multi-task to train the encoder in a supervised way. Then, we develop an attention based cross-modality temporal alignment method that combines the sequential video frames and transcript sentences. Finally, we adopt a copy mechanism to enable the decoder(generation) module to copy important concepts from source transcript directly. The extensive experimental results demonstrate the effectiveness of our model, which achieves state-of-the-art results on YouCookII dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eb3a42ce-4a37-48b7-a77a-6a727b169e38Cited by top-tier papers7
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng et al.ACM MM 2021 · 66 citations
- State-aware Video Procedural CaptioningTaichi Nishimura, Atsushi Hashimoto, Yoshitaka Ushiku, Hirotaka Kameko et al.ACM MM 2021 · 15 citations
- Narrative Action Evaluation with Prompt-Guided Multimodal InteractionShiyi Zhang, Sule Bai, Guangyi Chen, Lei Chen et al.CVPR 2024 · 12 citations
- MAMS: Model-Agnostic Module Selection Framework for Video CaptioningSangho Lee, Il Yong Chun, Hogun ParkAAAI 2025 · 1 citation
Related papers
- Text with Knowledge Graph Augmented Transformer for Video CaptioningXin Gu, Guang Chen, Yufei Wang, Libo Zhang et al.CVPR 2023
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi et al.CVPR 2024
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li et al.AAAI 2026 · 1 citation
- Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningYifan Lu, Ziqi Zhang, Chunfeng Yuan, Peng Li et al.AAAI 2024 · 7 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
