CSTA: CNN-based Spatiotemporal Attention for Video Summarization
Jaewon Son, Jaehun Park, Kwangsu Kim
Abstract
Video summarization aims to generate a concise repre-sentation of a video, capturing its essential content and key moments while reducing its overall length. Although several methods employ attention mechanisms to handle long-term dependencies, they often fail to capture the visual signif-icance inherent in frames. To address this limitation, we propose a CNN-based SpatioTemporal Attention (CSTA) method that stacks each feature of frames from a single video to form image-like frame representations and applies 2D CNN to these frame features. Our methodology relies on CNN to comprehend the inter and intra-frame relations and to find crucial attributes in videos by exploiting its abil-ity to learn absolute positions within images. In contrast to previous work compromising efficiency by designing additional modules to focus on spatial importance, CSTA re-quires minimal computational overhead as it uses CNN as a sliding window. Extensive experiments on two benchmark datasets (SumMe and TVSum) demonstrate that our pro-posed approach achieves state-of-the-art performance with fewer MACs compared to previous methods. Codes are available at https://github.com/thswodnjs3/CSTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9ad34db-03e2-4cfb-91c6-b3d715780a6cCited by top-tier papers7
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy et al.CVPR 2026 · 5 citations
- Video2BEV: Transforming Drone Videos to BEVs for Video-Based Geo-LocalizationHao Ju, Shaofei Huang, Si Liu, Zhedong ZhengICCV 2025 · 5 citations
- TripleSumm: Adaptive Triple-Modality Fusion for Video SummarizationSumin Kim, Hyemin Jeong, Mingu Kang, Yejin Kim et al.ICLR 2026 · 2 citations
- SummDiff: Generative Modeling of Video Summarization with DiffusionKwanseok Kim, Jaehoon Hahm, Sumin Kim, Jinhwan Sul et al.ICCV 2025 · 1 citation
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men et al.CVPR 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang et al.CVPR 2022 · 839 citations
- Incorporating Convolution Designs into Visual TransformersKun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou et al.ICCV 2021 · 581 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
Related papers
- VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal ModelingYamiao Ding, Tianrui Liu, Zhizhou Lu, Jun-Jie Huang et al.ACM MM 2025 · 1 citation
- Convolutional Hierarchical Attention Network for Query-Focused Video SummarizationShuwen Xiao, Zhou Zhao, Zijian Zhang, Xiaohui Yan et al.AAAI 2020 · 2 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin et al.CVPR 2020
