Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-Based Action Recognition
Wenhan Wu, Zhishuai Guo, Chen Chen, Hongfei Xue, Aidong Lu
Abstract
Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations but often overlooked the importance of finegrained action patterns in the semantic space (e.g., the hand movements in drinking water and brushing teeth). To address these limitations, we propose a Frequency-Semantic Enhanced Variational Autoencoder (FS-VAE) to explore the skeleton semantic representation learning with frequency decomposition. FS-VAE consists of three key components: 1) a frequency-based enhancement module with high-and low-frequency adjustments to enrich the skeletal semantics learning and improve the robustness of zero-shot action recognition; 2) a semantic-based action description with multilevel alignment to capture both local details and global correspondence, effectively bridging the semantic gap and compensating for the inherent loss of information in skeleton sequences; 3) a calibrated cross-alignment loss that enables valid skeleton-text pairs to counterbalance ambiguous ones, mitigating discrepancies and ambiguities in skeleton and text features, thereby ensuring robust alignment. Evaluations on the benchmarks demonstrate the effectiveness of our approach, validating that frequency-enhanced semantic features enable robust differentiation of visually and semantically similar action clusters, thereby improving zero-shot action recognition. Our project is publicly available at: https://github. com/wenhanwu95/FS-VAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- Learning Trajectory Dependencies for Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu Salzmann, Hongdong LiICCV 2019 · 534 citations
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee et al.CVPR 2022 · 383 citations
Related papers
- Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and MaximizationYujie Zhou, Wenwen Qiang, Anyi Rao, Ning Lin et al.ACM MM 2023 · 25 citations
- Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Tian He, Xiaocheng Lu et al.ACM MM 2024 · 13 citations
- Skeleton Motion Words for Unsupervised Skeleton-Based Temporal Action SegmentationUzay Gökay, Federico Spurio, Dominik R. Bach, Juergen GallICCV 2025 · 1 citation
- Generative Action Description Prompts for Skeleton-based Action RecognitionWangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang et al.ICCV 2023 · 84 citations
- Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed TransformerWenhan Wu, Ce Zheng, Zihao Yang, Chen Chen et al.ACM MM 2024 · 16 citations
