Use Your Head: Improving Long-Tail Video Recognition
Toby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi, Dima Damen
Abstract
This paper presents an investigation into long-tail video recognition. We demonstrate that, unlike naturallycollected video datasets and existing long-tail image benchmarks, current video benchmarks fall short on multiple long-tailed properties. Most critically, they lack few-shot classes in their tails. In response, we propose new video benchmarks that better assess long-tail recognition, by sampling subsets from two datasets: SSv2 and VideoLT. We then propose a method, Long-Tail Mixed Reconstruction (LMR), which reduces overfitting to instances from few-shot classes by reconstructing them as weighted combinations of samples from head classes. LMR then employs label mixing to learn robust decision boundaries. It achieves state-of-the-art average class accuracy on EPIC-KITCHENS and the proposed SSv2-LT and VideoLT-LT. Benchmarks and code at: github.com/ tobyperrett/lmr 1 We use the term 'naturally' to focus on the data collection. It does not imply footage of nature. We hope this footnote prevents any confusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5e8fab3-4232-46a7-a3c9-5e056f04d7e0Cited by top-tier papers6
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
- What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and LocationsChiara Plizzari, Toby Perrett, Barbara Caputo, Dima DamenICCV 2023 · 27 citations
- ELTA: An Enhancer against Long-Tail for Aesthetics-oriented ModelsLimin Liu, Shuai He, Anlong Ming, Rui Xie et al.ICML 2024 · 13 citations
- Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional PropertiesKeunwoo Peter Yu, Zheyuan Zhang, Fengyuan Hu, Shane Storks et al.EMNLP 2024 · 6 citations
- The PanAf-FGBG Dataset: Understanding the Impact of Backgrounds in Wildlife Behaviour RecognitionOtto Brookes, Maksim Kukushkin, Majid Mirmehdi, Colleen Stephens et al.CVPR 2025
Builds on41
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Balanced Meta-Softmax for Long-Tailed Visual RecognitionJiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma et al.NeurIPS 2020 · 861 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
Related papers
- VideoLT: Large-scale Long-tailed Video RecognitionXing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu et al.ICCV 2021 · 51 citations
- MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video RecognitionXinjie Li, Huijuan XuAAAI 2023 · 10 citations
- Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed RecognitionWonJun Moon, Hyun Seok Seong, Jae-Pil HeoAAAI 2023 · 6 citations
- Long-Tailed Anomaly Detection with Learnable Class NamesChih-Hui Ho, Kuan-Chuan Peng, Nuno VasconcelosCVPR 2024
- Exploring Long Tail Visual Relationship Recognition with Large VocabularySherif Abdelkarim, Aniket Agarwal, Panos Achlioptas, Jun Chen et al.ICCV 2021 · 19 citations
