Temporal Alignment-Free Video Matching for Few-shot Action Recognition
SuBeen Lee, WonJun Moon, Hyun Seok Seong, Jae-Pil Heo
Abstract
Few-Shot Action Recognition (FSAR) aims to train a model with only a few labeled video instances. A key challenge in FSAR is handling divergent narrative trajectories for precise video matching. While the frame-and tuple-level alignment approaches have been promising, their methods heavily rely on pre-defined and length-dependent alignment units (e.g., frames or tuples), which limits flexibility for actions of varying lengths and speeds. In this work, we introduce a novel TEmporal Alignment-free Matching (TEAM) approach, which eliminates the need for temporal units in action representation and brute-force alignment during matching. Specifically, TEAM represents each video with a fixed set of pattern tokens that capture globally discriminative clues within the video instance regardless of action length or speed, ensuring its flexibility. Furthermore, TEAM is inherently efficient, using token-wise comparisons to measure similarity between videos, unlike existing methods that rely on pairwise comparisons for temporal alignment. Additionally, we propose an adaptation process that identifies and removes common information across classes, establishing clear boundaries even between novel categories. Extensive experiments demonstrate the effectiveness of TEAM. Codes are available at github.com/leesb7426/TEAM. 'Up in the air' 'Getting ready to dive' 'Getting into the water' Token-wise Similarity Measurement 'Up in the air' 'Getting ready to dive' 'Getting into the water' ••• Frames Query Video Attention Degree Attention Degree
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4342ebb0-36df-4d13-9cff-24d5c11d65e5Cited by top-tier papers4
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang et al.ICCV 2025 · 4 citations
- Task-Specific Distance Correlation Matching for Few-Shot Action RecognitionFei Long, Yao Zhang, Jiaming Lv, Jiangtao Xie et al.AAAI 2026
- Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action RecognitionHantao Qi, Yan Yan, Junlong Gao, Hanzi WangCVPR 2026
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla et al.ICCV 2025
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li et al.AAAI 2022 · 98 citations
- Searching for Better Spatio-temporal Alignment in Few-Shot Action RecognitionYichao Cao, Xiu Su, Qingfei Tang, Shan You et al.NeurIPS 2022 · 13 citations
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu et al.CVPR 2022 · 73 citations
- Few-Shot Video Classification via Temporal AlignmentKaidi Cao, Jingwei Ji, Zhangjie Cao, Chien-Yi Chang et al.CVPR 2020
- Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionTianwei Yu, Peng Chen, Yuanjie Dang, Ruohong Huan et al.ACM MM 2023 · 10 citations
