METAL: Minimum Effort Temporal Activity Localization in Untrimmed Videos
Da Zhang, Xiyang Dai, Yuan-Fang Wang
Abstract
Existing Temporal Activity Localization (TAL) methods largely adopt strong supervision for model training which requires (1) vast amounts of untrimmed videos per each activity category and ( 2 ) accurate segment-level boundary annotations (start time and end time) for every instance. This poses a critical restriction to the current methods in practical scenarios where not only segment-level annotations are expensive to obtain but many activity categories are also rare and unobserved during training. Therefore, Can we learn a TAL model under weak supervision that can localize unseen activity classes? To address this scenario, we define a novel example-based TAL problem called Minimum Effort Temporal Activity Localization (METAL): Given only a few examples, the goal is to find the occurrences of semantically-related segments in an untrimmed video sequence while model training is only supervised by the video-level annotation. Towards this objective, we propose a novel Similarity Pyramid Network (SPN) that adopts the few-shot learning technique of Relation Network and directly encodes hierarchical multi-scale correlations, which we learn by optimizing two complimentary loss functions in an end-to-end manner. We evaluate the SPN on the THU-MOS'14 and ActivityNet datasets, of which we rearrange the videos to fit the METAL setup. Results show that our SPN achieves performance superior or competitive to stateof-the-art approaches with stronger supervision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web VideosTomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev et al.CVPR 2022 · 19 citations
- Annotation-Efficient Untrimmed Video Action RecognitionYixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian et al.ACM MM 2021 · 7 citations
- Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal DynamicsJuntae Lee, Mihir Jain, Sungrack YunICCV 2023 · 5 citations
- Reducing the Label Bias for Timestamp Supervised Temporal Action SegmentationKaiyuan Liu, Yunheng Li, Shenglan Liu, Chenwei Tan et al.CVPR 2023
- Few-Shot Transformation of Common Actions Into Time and SpacePengwan Yang, Pascal Mettes, Cees G. M. SnoekCVPR 2021
Related papers
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
- A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationAshraful Islam, Chengjiang Long, Richard J. RadkeAAAI 2021 · 145 citations
- Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksZiyi Liu, Le Wang, Qilin Zhang, Zhanning Gao et al.ICCV 2019 · 122 citations
- ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationZiyi Liu, Le Wang, Qilin Zhang, Wei Tang et al.AAAI 2021 · 83 citations
- Relational Prototypical Network for Weakly Supervised Temporal Action LocalizationLinjiang Huang, Yan Huang, Wanli Ouyang, Liang WangAAAI 2020 · 71 citations
