Temporal Alignment-Free Video Matching for Few-shot Action Recognition
SuBeen Lee, WonJun Moon, Hyun Seok Seong, Jae-Pil Heo
摘要
Few-Shot Action Recognition (FSAR) aims to train a model with only a few labeled video instances. A key challenge in FSAR is handling divergent narrative trajectories for precise video matching. While the frame-and tuple-level alignment approaches have been promising, their methods heavily rely on pre-defined and length-dependent alignment units (e.g., frames or tuples), which limits flexibility for actions of varying lengths and speeds. In this work, we introduce a novel TEmporal Alignment-free Matching (TEAM) approach, which eliminates the need for temporal units in action representation and brute-force alignment during matching. Specifically, TEAM represents each video with a fixed set of pattern tokens that capture globally discriminative clues within the video instance regardless of action length or speed, ensuring its flexibility. Furthermore, TEAM is inherently efficient, using token-wise comparisons to measure similarity between videos, unlike existing methods that rely on pairwise comparisons for temporal alignment. Additionally, we propose an adaptation process that identifies and removes common information across classes, establishing clear boundaries even between novel categories. Extensive experiments demonstrate the effectiveness of TEAM. Codes are available at github.com/leesb7426/TEAM. 'Up in the air' 'Getting ready to dive' 'Getting into the water' Token-wise Similarity Measurement 'Up in the air' 'Getting ready to dive' 'Getting into the water' ••• Frames Query Video Attention Degree Attention Degree
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang 等ICCV 2025 · 被引用 4 次
- Task-Specific Distance Correlation Matching for Few-Shot Action RecognitionFei Long, Yao Zhang, Jiaming Lv, Jiangtao Xie 等AAAI 2026
- Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action RecognitionHantao Qi, Yan Yan, Junlong Gao, Hanzi WangCVPR 2026
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla 等ICCV 2025
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li 等AAAI 2022 · 被引用 98 次
- Searching for Better Spatio-temporal Alignment in Few-Shot Action RecognitionYichao Cao, Xiu Su, Qingfei Tang, Shan You 等NeurIPS 2022 · 被引用 13 次
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu 等CVPR 2022 · 被引用 73 次
- Few-Shot Video Classification via Temporal AlignmentKaidi Cao, Jingwei Ji, Zhangjie Cao, Chien-Yi Chang 等CVPR 2020
- Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionTianwei Yu, Peng Chen, Yuanjie Dang, Ruohong Huan 等ACM MM 2023 · 被引用 10 次
