OnlineTAS: An Online Baseline for Temporal Action Segmentation
Qing Zhong, Guodong Ding, Angela Yao
摘要
Temporal context plays a significant role in temporal action segmentation. In an offline setting, the context is typically captured by the segmentation network after observing the entire sequence. However, capturing and using such context information in an online setting remains an under-explored problem. This work presents the an online framework for temporal action segmentation. At the core of the framework is an adaptive memory designed to accommodate dynamic changes in context over time, alongside a feature augmentation module that enhances the frames with the memory. In addition, we propose a post-processing approach to mitigate the severe over-segmentation in the online setting. On three common segmentation benchmarks, our approach achieves state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 被引用 6 次
- FlowNar: Scalable Streaming Narration for Long-Form VideosZeyun Zhong, Manuel Martin, Chengzhi Wu, David Schneider 等ICML 2026 · 被引用 1 次
- Condensing Action Segmentation Datasets via Generative Network InversionGuodong Ding, Rongyu Chen, Angela YaoCVPR 2025
- DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosZijia Lu, A S. M. Iftekhar, Gaurav Mittal, Tianjian Meng 等CVPR 2025
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Temporal Recurrent Networks for Online Action DetectionMingze Xu, Mingfei Gao, Yi-Ting Chen, Larry Davis 等ICCV 2019 · 被引用 201 次
- Long Short-Term Transformer for Online Action DetectionMingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li 等NeurIPS 2021 · 被引用 196 次
- MinVIS: A Minimal Video Instance Segmentation Framework without Video-based TrainingDe-An Huang, Zhiding Yu, Anima AnandkumarNeurIPS 2022 · 被引用 135 次
相关 Paper
- Hierarchical Event Memory for Accurate and Low-Latency Online Video Temporal GroundingMinghang Zheng, Yuxin Peng, Benyuan Sun, Yi Yang 等ICCV 2025 · 被引用 3 次
- MemorySeg: Online LiDAR Semantic Segmentation with a Latent MemoryEnxu Li, Sergio Casas, Raquel UrtasunICCV 2023 · 被引用 26 次
- Backtrace Mamba: Reviving Critical Temporal Contexts via Hierarchical Memory Compression for Online Action DetectionSu Yan, Jiahua Li, Kun Wei, Cheng DengAAAI 2026
- Coherent Temporal Synthesis for Incremental Action SegmentationGuodong Ding, Hans Golong, Angela YaoCVPR 2024
- Memory-based Adapters for Online 3D Scene PerceptionXiuwei Xu, Chong Xia, Ziwei Wang, Linqing Zhao 等CVPR 2024 · 被引用 4 次
