CSTA: CNN-based Spatiotemporal Attention for Video Summarization
Jaewon Son, Jaehun Park, Kwangsu Kim
摘要
Video summarization aims to generate a concise repre-sentation of a video, capturing its essential content and key moments while reducing its overall length. Although several methods employ attention mechanisms to handle long-term dependencies, they often fail to capture the visual signif-icance inherent in frames. To address this limitation, we propose a CNN-based SpatioTemporal Attention (CSTA) method that stacks each feature of frames from a single video to form image-like frame representations and applies 2D CNN to these frame features. Our methodology relies on CNN to comprehend the inter and intra-frame relations and to find crucial attributes in videos by exploiting its abil-ity to learn absolute positions within images. In contrast to previous work compromising efficiency by designing additional modules to focus on spatial importance, CSTA re-quires minimal computational overhead as it uses CNN as a sliding window. Extensive experiments on two benchmark datasets (SumMe and TVSum) demonstrate that our pro-posed approach achieves state-of-the-art performance with fewer MACs compared to previous methods. Codes are available at https://github.com/thswodnjs3/CSTA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy 等CVPR 2026 · 被引用 5 次
- Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-LocalizationHao Ju, Shaofei Huang, Si Liu, Zhedong ZhengICCV 2025 · 被引用 5 次
- TripleSumm: Adaptive Triple-Modality Fusion for Video SummarizationSumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim 等ICLR 2026 · 被引用 2 次
- SummDiff: Generative Modeling of Video Summarization with DiffusionKwanseok Kim, Jaehoon Hahm, Sumin Kim, Jinhwan Sul 等ICCV 2025 · 被引用 1 次
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men 等CVPR 2026
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang 等CVPR 2022 · 被引用 839 次
- Incorporating Convolution Designs into Visual TransformersKun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou 等ICCV 2021 · 被引用 581 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
相关 Paper
- VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal ModelingYamiao Ding, Tianrui Liu, Zhizhou Lu, Jun-Jie Huang 等ACM MM 2025 · 被引用 1 次
- Convolutional Hierarchical Attention Network for Query-Focused Video SummarizationShuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan 等AAAI 2020 · 被引用 2 次
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin 等CVPR 2020
