VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning
Kashu Yamazaki, Khoa Vo, Quang Sang Truong, Bhiksha Raj, Ngan Le
摘要
Video Paragraph Captioning aims to generate a multi-sentence description of an untrimmed video with multiple temporal event locations in a coherent storytelling. Following the human perception process, where the scene is effectively understood by decomposing it into visual (e.g. human, animal) and non-visual components (e.g. action, relations) under the mutual influence of vision and language, we first propose a visual-linguistic (VL) feature. In the proposed VL feature, the scene is modeled by three modalities including (i) a global visual environment; (ii) local visual main agents; (iii) linguistic scene elements. We then introduce an autoregressive Transformer-in-Transformer (TinT) to simultaneously capture the semantic coherence of intra- and inter-event contents within a video. Finally, we present a new VL contrastive loss function to guarantee the learnt embedding features are consistent with the captions semantics. Comprehensive experiments and extensive ablation studies on the ActivityNet Captions and YouCookII datasets show that the proposed Visual-Linguistic Transformer-in-Transform (VLTinT) outperforms previous state-of-the-art methods in terms of accuracy and diversity. The source code is made publicly available at: https://github.com/UARK-AICV/VLTinT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 FramesPinelopi Papalampidi, Skanda Koppula, Shreya Pathak, Justin Chiu 等CVPR 2024 · 被引用 15 次
- Models See Hallucinations: Evaluating the Factuality in Video CaptioningHui Liu, Xiaojun WanEMNLP 2023 · 被引用 5 次
- TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingYuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin 等AAAI 2026
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
相关 Paper
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
- Object Relation Attention for Image Paragraph CaptioningLi-Chuan Yang, Chih-Yuan Yang, Jane Yung-jen HsuAAAI 2021 · 被引用 17 次
- Towards Diverse Paragraph Captioning for Untrimmed VideosYuqing Song, Shizhe Chen, Qin JinCVPR 2021
- Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationWangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue 等CVPR 2022 · 被引用 23 次
- Text with Knowledge Graph Augmented Transformer for Video CaptioningXin Gu, Guang Chen, Yufei Wang, Libo Zhang 等CVPR 2023
