VideoLT: Large-scale Long-tailed Video Recognition
Xing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu, Jingjing Chen, Yu-Gang Jiang, Larry Davis
摘要
Label distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively studied for image classification tasks, limited effort has been made for the video domain. In this paper, we introduce VideoLT, a large-scale long-tailed video recognition dataset, as a step toward real-world video recognition. VideoLT contains 256,218 untrimmed videos, annotated into 1,004 classes with a long-tailed distribution. Through extensive studies, we demonstrate that state-of-the-art methods used for long-tailed image recognition do not perform well in the video domain due to the additional temporal dimension in videos. This motivates us to propose FrameStack, a simple yet effective method for long-tailed video recognition. In particular, FrameStack performs sampling at the frame-level in order to balance class distributions, and the sampling ratio is dynamically determined using knowledge derived from the network during training. Experimental results demonstrate that FrameStack can improve classification performance without sacrificing the overall accuracy. Code and dataset are available at: https://github.com/17Skye17/VideoLT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni 等CVPR 2022 · 被引用 102 次
- Discovering Objects that Can MoveZhipeng Bao, Pavel Tokmakov, Allan Jabri, Yu-Xiong Wang 等CVPR 2022 · 被引用 32 次
- Kill Two Birds with One Stone: Rethinking Data Augmentation for Deep Long-tailed LearningBinwu Wang, Pengkun Wang, Wei Xu, Xu Wang 等ICLR 2024 · 被引用 19 次
- Probability Guided Loss for Long-Tailed Multi-Label Image ClassificationDekun LinAAAI 2023 · 被引用 17 次
- Decoupled Optimisation for Long-Tailed Visual RecognitionCong Cong, Shiyu Xuan, Sidong Liu, Shiliang Zhang 等AAAI 2024 · 被引用 10 次
它引用的顶会 Paper12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 被引用 533 次
相关 Paper
- MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video RecognitionXinjie Li, Huijuan XuAAAI 2023 · 被引用 10 次
- Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed RecognitionWonJun Moon, Hyun Seok Seong, Jae-Pil HeoAAAI 2023 · 被引用 6 次
- Use Your Head: Improving Long-Tail Video RecognitionToby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi 等CVPR 2023
- Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxYu Li, Tao Wang, Bingyi Kang, Sheng Tang 等CVPR 2020
- Inflated Episodic Memory With Region Self-Attention for Long-Tailed Visual RecognitionLinchao Zhu, Yi YangCVPR 2020
