Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
Songping Wang, Qinglong Liu, Yueming Lyu, Ning Li, Ziwen He, Caifeng Shan
摘要
Mixture-of-Experts (MoE) has demonstrated strong performance in video understanding tasks, yet its adversarial robustness remains underexplored. Existing attack methods often treat MoE as a unified architecture, overlooking the independent and collaborative weaknesses of key components such as routers and expert modules. To fill this gap, we propose Temporal Lipschitz-Guided Attacks (TLGA) to thoroughly investigate component-level vulnerabilities in video MoE models. We first design attacks on the router, revealing its independent weaknesses. Building on this, we introduce Joint Temporal Lipschitz-Guided Attacks (J-TLGA), which collaboratively perturb both routers and experts. This joint attack significantly amplifies adversarial effects and exposes the Achilles’ Heel (collaborative weaknesses) of the MoE architecture. Based on these insights, we further propose Joint Temporal Lipschitz Adversarial Training (J-TLAT). J-TLAT performs joint training to further defend against collaborative weaknesses, enhancing component-wise robustness. Our framework is plug-and-play and reduces inference cost by more than 60% compared with dense models. It consistently enhances adversarial robustness across diverse video datasets and model architectures, effectively mitigating both the independent and collaborative weaknesses of MoE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang 等ICLR 2026 · 被引用 44 次
- LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency ExpertsChen Zhao, Jiawei Chen, Hongyu Li, Zhuoliang Kang 等ICML 2026 · 被引用 16 次
- Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object DetectionHuafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng ShanCVPR 2026
它引用的顶会 Paper22
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Rethinking Lipschitz Neural Networks and Certified Robustness: A Boolean Function PerspectiveBohang Zhang, Du Jiang, Di He, Liwei WangNeurIPS 2022 · 被引用 88 次
相关 Paper
- Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model ApproachXu Zhang, Kaidi Xu, Ziqing Hu, Ren WangICML 2025
- Robust Mixture-of-Expert Training for Convolutional Neural NetworksYihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang 等ICCV 2023 · 被引用 43 次
- Timeexpert: an Expert-Guided Video Llm for Video Temporal GroundingZuhao Yang, Yingchen Yu, Yunqing Zhao, Shijian Lu 等ICCV 2025 · 被引用 3 次
- SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignmentQingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks 等CCS 2026
- ReMoE: Region-Mixture Experts for Adversarially-Robust Vision TransformersQinghao Zhong, Bingzhi Chen, Yishu Liu, Minhua Lu 等CVPR 2026
