FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual Compression
Bo Tong, Bokai Lai, Yiyi Zhou, Gen Luo, Yunhang Shen, Ke Li, Xiaoshuai Sun, Rongrong Ji
Abstract
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their actual speedup. In this paper, we propose a powerful and fast tiny MLLM called FlashSloth. Different from previous efforts, FlashSloth focuses on improving the descriptive power of visual tokens in the process of compressing their redundant semantics. In particular, Flash-Sloth introduces embedded visual compression designs to capture both visually salient and instruction-related image information, so as to achieving superior multimodal performance with fewer visual tokens. Extensive experiments are conducted to validate the proposed FlashSloth, and a bunch of tiny but strong MLLMs are also comprehensively compared, e.g., InternVL2, MiniCPM-V2 and Qwen2-VL. The experimental results show that compared with these advanced tiny MLLMs, our FlashSloth can greatly reduce the number of visual tokens, training memory and computation complexity while retaining high performance on various VL tasks. Our code is released at: https://github.com/ codefanw/FlashSloth.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22dc8745-d04f-44a5-be32-aee81207d943Cited by top-tier papers4
- Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory MechanismTao Chen, Kun Zhang, Qiong Wu, Xiao Chen et al.CVPR 2026 · 8 citations
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 6 citations
- A More Word-like Image Tokenization for MLLMsHyun Lee, Hyemin Jeong, Yejin Kim, Hyungwook Choi et al.CVPR 2026 · 2 citations
- Efficient Multimodal Spatial Reasoning via Dynamic and Asymmetric RoutingYixian Shen, Qi Bi, Zihan Wang, Zhiheng Yang et al.ICLR 2026
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang et al.ICLR 2026 · 3 citations
- ApET: Approximation-Error Guided Token Compression for Efficient VLMsQiankun Ma, Ziyao Zhang, Haofei Wang, Zhen Song et al.CVPR 2026 · 11 citations
- Accelerating Multimodal Large Language Models by Searching Optimal Vision Token ReductionShiyu Zhao, Zhenting Wang, Felix Juefei-Xu, Xide Xia et al.CVPR 2025
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token SkippingWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao YanICCV 2025
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye et al.NeurIPS 2025 · 16 citations
