Alignment-Uniformity aware Representation Learning for Zero-shot Video Classification
Shi Pu, Kaili Zhao, Mao Zheng
Abstract
Most methods tackle zero-shot video classification by aligning visual-semantic representations within seen classes, which limits generalization to unseen classes. To enhance model generalizability, this paper presents an endto-end framework that preserves alignment and uniformity properties for representations on both seen and unseen classes. Specifically, we formulate a supervised contrastive loss to simultaneously align visual-semantic features (i.e., alignment) and encourage the learned features to distribute uniformly (i.e., uniformity). Unlike existing methods that only consider the alignment, we propose uniformity to preserve maximal-info of existing features, which improves the probability that unobserved features fall around observed data. Further, we synthesize features of unseen classes by proposing a class generator that interpolates and extrapolates the features of seen classes. Besides, we introduce two metrics, closeness and dispersion, to quantify the two properties and serve as new measurements of model generalizability. Experiments show that our method significantly outperforms SoTA by relative improvements of 28.1% on UCF101 and 27.0% on HMDB51. Code is available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Rethinking Graph Masked Autoencoders through Alignment and UniformityLiang Wang, Xiang Tao, Qiang Liu, Shu Wu et al.AAAI 2024 · 40 citations
- Generalizable Person Re-identification via Balancing Alignment and UniformityYoonki Cho, Jaeyoon Kim, Woo Jae Kim, Junsik Jung et al.NeurIPS 2024 · 21 citations
- Semantic Evolvement Enhanced Graph Autoencoder for Rumor DetectionXiang Tao, Liang Wang, Qiang Liu, Shu Wu et al.WWW 2024 · 20 citations
- Your contrastive learning problem is secretly a distribution alignment problemZihao Chen, Chi-Heng Lin, Ran Liu, Jingyun Xiao et al.NeurIPS 2024 · 14 citations
- ReGen: A good Generative zero-shot video classifier should be RewardedAdrian Bulat, Enrique Sanchez, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 1 citation
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Semantics Disentangling for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang et al.ICCV 2021 · 143 citations
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
- Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- FREE: Feature Refinement for Generalized Zero-Shot LearningShiming Chen, Wenjie Wang, Beihao Xia, Qinmu Peng et al.ICCV 2021 · 171 citations
- Generalized Zero-Shot Learning using Generated Proxy Unseen Samples and Entropy SeparationOmkar Gune, Biplab Banerjee, Subhasis Chaudhuri, Fabio CuzzolinACM MM 2020 · 15 citations
