How Much Temporal Long-Term Context is Needed for Action Segmentation?
Emad Bahrami Rad, Gianpiero Francesca, Juergen Gall
Abstract
Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal performance. While transformers can model the long-term context of a video, this becomes computationally prohibitive for long videos. Recent works on temporal action segmentation thus combine temporal convolutional networks with self-attentions that are computed only for a local temporal window. While these approaches show good results, their performance is limited by their inability to capture the full context of a video. In this work, we try to answer how much long-term temporal context is required for temporal action segmentation by introducing a transformer-based model that leverages sparse attention to capture the full context of a video. We compare our model with the current state of the art on three datasets for temporal action segmentation, namely 50Salads, Breakfast, and Assembly101. Our experiments show that modeling the full context of a video is necessary to obtain the best performance for temporal action segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8660d2b4-b1ba-4b40-bac7-8ca1729a832bCited by top-tier papers16
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Efficient Temporal Action Segmentation via Boundary-aware Query VotingPeiyao Wang, Yuewei Lin, Erik Blasch, Jie Wei et al.NeurIPS 2024 · 30 citations
- Hierarchical Vector Quantization for Unsupervised Action SegmentationFederico Spurio, Emad Bahrami, Gianpiero Francesca, Juergen GallAAAI 2025 · 17 citations
- OnlineTAS: An Online Baseline for Temporal Action SegmentationQing Zhong, Guodong Ding, Angela YaoNeurIPS 2024 · 15 citations
- ActFusion: a Unified Diffusion Model for Action Segmentation and AnticipationDayoung Gong, Suha Kwak, Minsu ChoNeurIPS 2024 · 14 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi et al.CVPR 2021
- Shifted Chunk Transformer for Spatio-Temporal Representational LearningXuefan Zha, Wentao Zhu, Xun Lv, Sen Yang et al.NeurIPS 2021 · 32 citations
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha et al.CVPR 2022 · 56 citations
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li et al.NeurIPS 2021 · 196 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
