MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video Recognition
Xinjie Li, Huijuan Xu
Abstract
The long-tailed video recognition problem is especially challenging, as videos tend to be long and untrimmed, and each video may contain multiple classes, causing frame-level class imbalance. The previous method tackles the long-tailed video recognition only through frame-level sampling for class re-balance without distinguishing the frame-level feature representation between head and tail classes. To improve the frame-level feature representation of tail classes, we modulate the frame-level features with an auxiliary distillation loss to reduce the distribution distance between head and tail classes. Moreover, we design a mixture-of-experts framework with two different expert designs, i.e., the first expert with an attention-based classification network handling the original long-tailed distribution, and the second expert dealing with the re-balanced distribution from class-balanced sampling. Notably, in the second expert, we specifically focus on the frames unsolved by the first expert through designing a complementary frame selection module, which inherits the attention weights from the first expert and selects frames with low attention weights, and we also enhance the motion feature representation for these selected frames. To highlight the multi-label challenge in long-tailed video recognition, we create two additional benchmarks based on Charades and CharadesEgo videos with the multi-label property, called CharadesLT and CharadesEgoLT. Extensive experiments are conducted on the existing long-tailed video benchmark VideoLT and the two new benchmarks to verify the effectiveness of our proposed method with state-of-the-art performance. The code and proposed benchmarks are released at https://github.com/VisionLanguageLab/MEID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Language-driven All-in-one Adverse Weather RemovalHao Yang, Liyuan Pan, Yan Yang, Wei LiangCVPR 2024 · 28 citations
- Mixture of Experts as Representation Learner for Deep Multi-View ClusteringYunhe Zhang, Jinyu Cai, Zhihao Wu, Pengyang Wang et al.AAAI 2025 · 16 citations
- GAMMA: Graph Neural Network-Based Multi-Bottleneck Localization for Microservices ApplicationsGagan Somashekar, Anurag Dutt, Mainak Adak, Tania Lorido-Botran et al.WWW 2024 · 15 citations
- Generative Active Learning for Long-Tail Trajectory Prediction via Controllable Diffusion ModelDaehee Park, Monu Surana, Pranav Desai, Ashish Mehta et al.ICCV 2025 · 2 citations
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
Builds on19
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Parametric Contrastive LearningJiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu et al.ICCV 2021 · 375 citations
Related papers
- VideoLT: Large-scale Long-tailed Video RecognitionXing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu et al.ICCV 2021 · 51 citations
- Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed RecognitionWonJun Moon, Hyun Seok Seong, Jae-Pil HeoAAAI 2023 · 6 citations
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object DetectionHao Vo, Khoa Vo, Thinh Phan, Ngo Xuan Cuong et al.CVPR 2026 · 1 citation
- Use Your Head: Improving Long-Tail Video RecognitionToby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2023
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 1 citation
