Make Your Training Flexible: Towards Deployment-Efficient Video Models
Chenting Wang, Kunchang Li, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, Limin Wang
Abstract
Current video training methods rely on fixed spatiotemporal sampling grids to extract a predetermined number of tokens, limiting adaptability to diverse computational budgets and resulting in suboptimal accuracy-computation trade-offs. This rigidity constrains high-performance models trained in resource-rich environments from being efficiently deployed on resource-constrained devices. We hence introduce a novel paradigm for lossless adaptation across scenarios, enabling models to maintain optimal performance under high-resource conditions while seamlessly transferring to low-resource environments. Central to this is Token Optimization (TO), an adaptive inference framework that namically samples and selects input token set to optimize input information under varied computational constraints. To support this, we propose Flux, an augmentation tool that enables flexible sampling grids and token selection. It integrates seamlessly into popular video training frameworks, significantly enhancing model robustness and adaptability with negligible additional cost. Applied to large-scale video pretraining, our method produces FluxViT, which achieves state-of-the-art performance across multiple tasks under standard costs. Remarkably, with only of the tokens, FluxViT matches prior state-of-the-art models under TO across tasks, achieving nearly 90% computational savings. Code and models are available at https://github.com/OpenGVLab/FluxViT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 359c3d4a-24b9-4a6c-a030-516f2bba130dCited by top-tier papers2
- SpecDiff: Accelerating Diffusion Model Inference with Self-SpeculationJiayi Pan, Jiaming Xu, Yongkang Zhou, Guohao DaiAAAI 2026 · 1 citation
- InternVideo-Next: Towards World-Understanding Video ModelsChenting Wang, Yuhan Zhu, Yicheng Xu, Jiange Yang et al.CVPR 2026
Builds on48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
Related papers
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video TokenizationAndrei Atanov, Jesse Allardice, Roman Bachmann, Oğuzhan Fatih Kar et al.ICML 2026 · 3 citations
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 10 citations
- Principles of Visual Tokens for Efficient Video UnderstandingXinyue Hao, Gen Li, Shreyank N. Gowda, Robert B. Fisher et al.ICCV 2025 · 1 citation
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu et al.ICML 2024 · 94 citations
- Learning Generalized Trackers with Elastic Token BudgetsYinchao Ma, Jianpeng Yang, Yuyang Tang, Jie Xiao et al.ICML 2026
