Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition
Bozheng Li, Mushui Liu, Gaoang Wang, Yunlong Yu
摘要
In this paper, we propose a novel Temporal Sequence-Aware-Model (TSAM) for few-shot action recognition (FSAR), which incorporates a sequential perceiver adapter into the pre-training framework, to integrate both the spatial information and the sequential temporal dynamics into the feature embeddings. Different from the existing fine-tuning approaches that capture temporal information by exploring the relationships among all the frames, our perceiver-based adapter recurrently captures the sequential dynamics alongside the timeline, which could perceive the frame order change. To obtain the discriminative representations for each class, we extend a textual corpus for each class derived from the large language models (LLMs) and enrich the visual prototypes by integrating the contextual semantic information. Besides, We introduce an unbalanced optimal transport strategy for feature matching that mitigates the impact of class-unrelated features, thereby facilitating more effective decision-making. Experimental results on five FSAR datasets demonstrate that our method establishes a new benchmark, outperforming the second-best competitors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Attentive Eraser: Unleashing Diffusion Model's Object Removal Potential via Self-Attention Redirection GuidanceWenhao Sun, Xue-Mei Dong, Benlei Cui, Jingqun TangAAAI 2025 · 被引用 50 次
- TFGDA: Exploring Topology and Feature Alignment in Semi-supervised Graph Domain Adaptation through Robust ClusteringJun Dan, Weiming Liu, Chunfeng Xie, Hua Yu 等NeurIPS 2024 · 被引用 22 次
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang 等ICCV 2025 · 被引用 4 次
- Video Repurposing from User Generated Content: A Large-scale Dataset and BenchmarkYongliang Wu, Wenbo Zhu, Jiawang Cao, Yi Lu 等AAAI 2025 · 被引用 2 次
- Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal GroundingZelin Zheng, Xinyan Liu, Ruixin Li, Antoni B. Chan 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper14
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang 等CVPR 2022 · 被引用 124 次
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu 等ACM MM 2020 · 被引用 97 次
相关 Paper
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv 等ACM MM 2024 · 被引用 11 次
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action RecognitionHuabin Liu, Weixian Lv, John See, Weiyao LinACM MM 2022 · 被引用 11 次
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi 等ACM MM 2023 · 被引用 20 次
- TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action RecognitionYilong Wang, Zilin Gao, Qilong Wang, Zhaofeng Chen 等CVPR 2025
