Alignment-Uniformity aware Representation Learning for Zero-shot Video Classification
Shi Pu, Kaili Zhao, Mao Zheng
摘要
Most methods tackle zero-shot video classification by aligning visual-semantic representations within seen classes, which limits generalization to unseen classes. To enhance model generalizability, this paper presents an endto-end framework that preserves alignment and uniformity properties for representations on both seen and unseen classes. Specifically, we formulate a supervised contrastive loss to simultaneously align visual-semantic features (i.e., alignment) and encourage the learned features to distribute uniformly (i.e., uniformity). Unlike existing methods that only consider the alignment, we propose uniformity to preserve maximal-info of existing features, which improves the probability that unobserved features fall around observed data. Further, we synthesize features of unseen classes by proposing a class generator that interpolates and extrapolates the features of seen classes. Besides, we introduce two metrics, closeness and dispersion, to quantify the two properties and serve as new measurements of model generalizability. Experiments show that our method significantly outperforms SoTA by relative improvements of 28.1% on UCF101 and 27.0% on HMDB51. Code is available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Rethinking Graph Masked Autoencoders through Alignment and UniformityLiang Wang, Xiang Tao, Qiang Liu, Shu Wu 等AAAI 2024 · 被引用 40 次
- Generalizable Person Re-identification via Balancing Alignment and UniformityYoonki Cho, Jaeyoon Kim, Woo Jae Kim, Junsik Jung 等NeurIPS 2024 · 被引用 21 次
- Semantic Evolvement Enhanced Graph Autoencoder for Rumor DetectionXiang Tao, Liang Wang, Qiang Liu, Shu Wu 等WWW 2024 · 被引用 20 次
- Your contrastive learning problem is secretly a distribution alignment problemZihao Chen, Chi-Heng Lin, Ran Liu, Jingyun Xiao 等NeurIPS 2024 · 被引用 14 次
- ReGen: A good Generative zero-shot video classifier should be RewardedAdrian Bulat, Enrique Sanchez, Brais Martínez, Georgios TzimiropoulosICCV 2023 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Semantics Disentangling for Generalized Zero-Shot LearningZhi Chen, Yadan Luo, Ruihong Qiu, Sen Wang 等ICCV 2021 · 被引用 143 次
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 被引用 13 次
- Primitive Generation and Semantic-Related Alignment for Universal Zero-Shot SegmentationShuting He, Henghui Ding, Wei JiangCVPR 2023
- FREE: Feature Refinement for Generalized Zero-Shot LearningShiming Chen, Wenjie Wang, Beihao Xia, Qinmu Peng 等ICCV 2021 · 被引用 171 次
- Generalized Zero-Shot Learning using Generated Proxy Unseen Samples and Entropy SeparationOmkar Gune, Biplab Banerjee, Subhasis Chaudhuri, Fabio CuzzolinACM MM 2020 · 被引用 15 次
