Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action Recognition
Pulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla, Abhinav Shrivastava
Abstract
Video understanding requires effective modeling of both motion and appearance information, particularly for few-shot action recognition. While recent advances in point tracking have been shown to improve few-shot action recognition, two fundamental challenges persist: selecting informative points to track and effectively modeling their motion patterns. We present Trokens, a novel approach that transforms trajectory points into semantic-aware relational tokens for action recognition. First, we introduce a semantic-aware sampling strategy to adaptively distribute tracking points based on object scale and semantic relevance. Second, we develop a motion modeling framework that captures both intra-trajectory dynamics through the Histogram of Oriented Displacements (HoD) and inter-trajectory relationships to model complex action patterns. Our approach effectively combines these trajectory tokens with semantic features to enhance appearance features with motion information, achieving state-of-the-art performance across six diverse few-shot action recognition benchmarks: Something-Something-V2 (both full and small splits), Kinetics, UCF101, HMDB51, and FineGym. For project page see https://trokens-iccv25.github.io
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bad0774b-8049-4fe6-bec3-ecc1ed6b3e58Cited by top-tier papers3
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment VideosSeungjae Lee, Yoonkyo Jung, Inkook Chun, Yao-Chih Lee et al.CVPR 2026 · 17 citations
- Frame2Freq: Spectral Adapters for Fine-Grained Video UnderstandingThinesh Thiyakesan Ponbagavathi, Constantin Seibold, Alina RoitbergCVPR 2026 · 2 citations
- Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action RecognitionHantao Qi, Yan Yan, Junlong Gao, Hanzi WangCVPR 2026
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Unsupervised Semantic Segmentation by Distilling Feature CorrespondencesMark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely et al.ICLR 2022 · 317 citations
- TAPIR: Tracking Any Point with per-frame Initialization and temporal RefinementCarl Doersch, Yi Yang, Mel Vecerík, Dilara Gokay et al.ICCV 2023 · 297 citations
Related papers
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2021
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu et al.ACM MM 2024 · 8 citations
- Revisiting the Spatial and Temporal Modeling for Few-Shot Action RecognitionJiazheng Xing, Mengmeng Wang, Yong Liu, Boyu MuAAAI 2023 · 51 citations
- Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionShuo Zheng, Yuanjie Dang, Peng Chen, Ruohong Huan et al.ACM MM 2024 · 1 citation
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang et al.CVPR 2022 · 124 citations
