Few-Shot Transformation of Common Actions Into Time and Space
Pengwan Yang, Pascal Mettes, Cees G. M. Snoek
Abstract
This paper introduces the task of few-shot common action localization in time and space. Given a few trimmed support videos containing the same but unknown action, we strive for spatio-temporal localization of that action in a long untrimmed query video. We do not require any class labels, interval bounds, or bounding boxes. To address this challenging task, we introduce a novel few-shot transformer architecture with a dedicated encoder-decoder structure optimized for joint commonality learning and localization prediction, without the need for proposals. Experiments on reorganizations of the AVA and UCF101-24 datasets show the effectiveness of our approach for few-shot common action localization, even when the support videos are noisy. Although we are not specifically designed for common localization in time only, we also compare favorably against the fewshot and one-shot state-of-the-art in this setting. Lastly, we demonstrate that the few-shot transformer is easily extended to common action localization per pixel.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af4a12be-350d-4b05-b0a7-638e1e0f81f8Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- SILCO: Show a Few Images, Localize the Common ObjectTao Hu, Pascal Mettes, Jia-Hong Huang, Cees SnoekICCV 2019 · 49 citations
- METAL: Minimum Effort Temporal Activity Localization in Untrimmed VideosDa Zhang, Xiyang Dai, Yuan-Fang WangCVPR 2020
Related papers
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
- Searching for Better Spatio-temporal Alignment in Few-Shot Action RecognitionYichao Cao, Xiu Su, Qingfei Tang, Shan You et al.NeurIPS 2022 · 13 citations
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2021
- TxVAD: Improved Video Action Detection by TransformersZhenyu Wu, Zhou Ren, Yi Wu, Zhangyang Wang et al.ACM MM 2022 · 5 citations
- QDETRv: Query-Guided DETR for One-Shot Object Localization in VideosYogesh Kumar, Saswat Mallick, Anand Mishra, Sowmya Rasipuram et al.AAAI 2024 · 4 citations
