Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Gagan Jain, Nidhi Hegde, Aditya Kusupati, Arsha Nagrani, Shyamal Buch, Prateek Jain, Anurag Arnab, Sujoy Paul
Abstract
The visual medium (images and videos) naturally contains a large amount of information redundancy, thereby providing a great opportunity for leveraging efficiency in processing. While Vision Transformer (ViT) based models scale effectively to large data regimes, they fail to capitalize on this inherent redundancy, leading to higher computational costs. Mixture of Experts (MoE) networks demonstrate scalability while maintaining same inference-time costs, but they come with a larger parameter footprint. We present Mixture of Nested Experts (MoNE), which utilizes a nested structure for experts, wherein individual experts fall on an increasing compute-accuracy curve. Given a compute budget, MoNE learns to dynamically choose tokens in a priority order, and thus redundant tokens are processed through cheaper nested experts. Using this framework, we achieve equivalent performance as the baseline models, while reducing inference time compute by over two-fold. We validate our approach on standard image and video datasets - ImageNet-21K, Kinetics400, and Something-Something-v2. We further highlight MoNEs adaptability by showcasing its ability to maintain strong performance across different inference-time compute budgets on videos, using only a single trained model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Advancing Expert Specialization for Better MoEHongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu et al.NeurIPS 2025 · 48 citations
- Sherlock: Towards Multi-scene Video Abnormal Event Extraction and Localization via a Global-local Spatial-sensitive LLMJunxiao Ma, Jingjing Wang, Jiamin Luo, Peiying Yu et al.WWW 2025 · 11 citations
- Adaptive Computation Modules: Granular Conditional Computation for Efficient InferenceBartosz Wójcik, Alessio Devoto, Karol Pustelnik, Pasquale Minervini et al.AAAI 2025 · 8 citations
- ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic InferenceAli Hojjat, Janek Haberer, Sören Pirk, Olaf LandsiedelCVPR 2026 · 6 citations
- Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly DetectionGiacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni et al.ICCV 2025 · 5 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Principles of Visual Tokens for Efficient Video UnderstandingXinyue Hao, Gen Li, Shreyank N. Gowda, Robert B. Fisher et al.ICCV 2025 · 1 citation
- Eventful Transformers: Leveraging Temporal Redundancy in Vision TransformersMatthew Dutson, Yin Li, Mohit GuptaICCV 2023 · 18 citations
- From Sparse to Soft Mixtures of ExpertsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Neil HoulsbyICLR 2024 · 264 citations
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech RecognitionUmberto Cappellazzo, Minsu Kim, Pingchuan Ma, Honglie Chen et al.NeurIPS 2025 · 5 citations
