MEID: Mixture-of-Experts with Internal Distillation for Long-Tailed Video Recognition
Xinjie Li, Huijuan Xu
摘要
The long-tailed video recognition problem is especially challenging, as videos tend to be long and untrimmed, and each video may contain multiple classes, causing frame-level class imbalance. The previous method tackles the long-tailed video recognition only through frame-level sampling for class re-balance without distinguishing the frame-level feature representation between head and tail classes. To improve the frame-level feature representation of tail classes, we modulate the frame-level features with an auxiliary distillation loss to reduce the distribution distance between head and tail classes. Moreover, we design a mixture-of-experts framework with two different expert designs, i.e., the first expert with an attention-based classification network handling the original long-tailed distribution, and the second expert dealing with the re-balanced distribution from class-balanced sampling. Notably, in the second expert, we specifically focus on the frames unsolved by the first expert through designing a complementary frame selection module, which inherits the attention weights from the first expert and selects frames with low attention weights, and we also enhance the motion feature representation for these selected frames. To highlight the multi-label challenge in long-tailed video recognition, we create two additional benchmarks based on Charades and CharadesEgo videos with the multi-label property, called CharadesLT and CharadesEgoLT. Extensive experiments are conducted on the existing long-tailed video benchmark VideoLT and the two new benchmarks to verify the effectiveness of our proposed method with state-of-the-art performance. The code and proposed benchmarks are released at https://github.com/VisionLanguageLab/MEID.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Language-driven All-in-one Adverse Weather RemovalHao Yang, Liyuan Pan, Yan Yang, Wei LiangCVPR 2024 · 被引用 28 次
- Mixture of Experts as Representation Learner for Deep Multi-View ClusteringYunhe Zhang, Jinyu Cai, Zhihao Wu, Pengyang Wang 等AAAI 2025 · 被引用 16 次
- GAMMA: Graph Neural Network-Based Multi-Bottleneck Localization for Microservices ApplicationsGagan Somashekar, Anurag Dutt, Mainak Adak, Tania Lorido-Botran 等WWW 2024 · 被引用 15 次
- Generative Active Learning for Long-Tail Trajectory Prediction via Controllable Diffusion ModelDaehee Park, Monu Surana, Pranav Desai, Ashish Mehta 等ICCV 2025 · 被引用 2 次
- Progressive Mask Distillation for Self-supervised Video RepresentationKewei Wu, Chong Liang, Zhao Xie, Dan GuoCVPR 2026
它引用的顶会 Paper19
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
- Parametric Contrastive LearningJiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu 等ICCV 2021 · 被引用 375 次
相关 Paper
- VideoLT: Large-scale Long-tailed Video RecognitionXing Zhang, Zuxuan Wu, Zejia Weng, Huazhu Fu 等ICCV 2021 · 被引用 51 次
- Minority-Oriented Vicinity Expansion with Attentive Aggregation for Video Long-Tailed RecognitionWonJun Moon, Hyun Seok Seong, Jae-Pil HeoAAAI 2023 · 被引用 6 次
- SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object DetectionHao Vo, Khoa Vo, Thinh Phan, Ngo Xuan Cuong 等CVPR 2026 · 被引用 1 次
- Use Your Head: Improving Long-Tail Video RecognitionToby Perrett, Saptarshi Sinha, Tilo Burghardt, Majid Mirmehdi 等CVPR 2023
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 被引用 1 次
