InsAT: Instance-aware Semantic Alignment and Transfer from Human-Object Keypoints for Zero-to-Few-shot Action Understanding
Kazuki Tsutsukawa
摘要
Keypoint-based action recognition offers robustness to appearance variations and provides privacy-preserving representation. However, existing zero-shot (ZS) approaches largely emphasize human motion while underutilizing contextual information, particularly human-object interactions. Moreover, extending keypoint-based ZS models to few-shot scenarios remains insufficiently explored. We propose Instance-aware Semantic Alignment and Transfer (InsAT), a unified framework for ZS recognition and zero-to-few-shot (Z2F) adaptation that leverages instance-level language descriptions. InsAT aligns textual descriptions of humans, objects, and their interactions with visual representations derived from human and object keypoints, enabling effective transfer of interaction knowledge from seen to unseen action classes. To support Z2F adaptation, we introduce Instance-level Visual Adaptation, a parameter-free mechanism that improves recognition by incorporating instance-level contextual cues without updating model weights. Extensive experiments demonstrate that InsAT substantially outperforms prior keypoint-based ZS methods and achieves competitive performance relative to large vision-language models, while remaining data-efficient and robust.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko 等EMNLP 2021 · 被引用 399 次
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 被引用 185 次
相关 Paper
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu 等NeurIPS 2024 · 被引用 45 次
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeWei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger 等ICCV 2023 · 被引用 52 次
- CIA: Class- and Instance-aware Adaptation for Vision-Language ModelsLin Peng, Cong Wan, Shaokun Wang, Xiang Song 等ACM MM 2025 · 被引用 1 次
- EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionQinqian Lei, Bo Wang, Robby T. TanNeurIPS 2024 · 被引用 42 次
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd 等CVPR 2026 · 被引用 3 次
