Use Your Head: Improving Long-Tail Video Recognition
Toby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi, Dima Damen
摘要
This paper presents an investigation into long-tail video recognition. We demonstrate that, unlike naturallycollected video datasets and existing long-tail image benchmarks, current video benchmarks fall short on multiple long-tailed properties. Most critically, they lack few-shot classes in their tails. In response, we propose new video benchmarks that better assess long-tail recognition, by sampling subsets from two datasets: SSv2 and VideoLT. We then propose a method, Long-Tail Mixed Reconstruction (LMR), which reduces overfitting to instances from few-shot classes by reconstructing them as weighted combinations of samples from head classes. LMR then employs label mixing to learn robust decision boundaries. It achieves state-of-the-art average class accuracy on EPIC-KITCHENS and the proposed SSv2-LT and VideoLT-LT. Benchmarks and code at: github.com/ tobyperrett/lmr 1 We use the term 'naturally' to focus on the data collection. It does not imply footage of nature. We hope this footnote prevents any confusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 被引用 28 次
- What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and LocationsChiara Plizzari, Toby Perrett, Barbara Caputo, Dima DamenICCV 2023 · 被引用 27 次
- ELTA: An Enhancer against Long-Tail for Aesthetics-oriented ModelsLimin Liu, Shuai He, Anlong Ming, Rui Xie 等ICML 2024 · 被引用 13 次
- Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional PropertiesKeunwoo Peter Yu, Zheyuan Zhang, Fengyuan Hu, Shane Storks 等EMNLP 2024 · 被引用 6 次
- The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionOtto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens 等CVPR 2025
它引用的顶会 Paper41
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma 等NeurIPS 2020 · 被引用 861 次
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin 等ICLR 2020 · 被引用 692 次
相关 Paper
- VideoLT: Large-scale Long-tailed Video RecognitionXing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu 等ICCV 2021 · 被引用 51 次
- MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video RecognitionXinjie Li, Huijuan XuAAAI 2023 · 被引用 10 次
- Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed RecognitionWonJun Moon, Hyun Seok Seong, Jae-Pil HeoAAAI 2023 · 被引用 6 次
- Long-Tailed Anomaly Detection with Learnable Class NamesChih-Hui Ho, Kuan-Chuan Peng, Nuno VasconcelosCVPR 2024
- Exploring Long Tail Visual Relationship Recognition with Large VocabularySherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen 等ICCV 2021 · 被引用 19 次
