OmniFit: Bridging Modalities via Layer-Adaptive Token Compression for Omnimodal Large Language Models
Zining Wang, Zhihang Yuan, Yingjie Zhai, Wenshuo Li, Han Shu, Ruihao Gong, Jinyang Guo, Xianglong Liu
Abstract
Emerging Omni-modal Large Language Models (OmniLLMs) enable real-time interaction across video, audio, and text but suffer from prohibitive computational costs due to the quadratic complexity of processing continuous streaming inputs. Existing token compression strategies remain suboptimal as they typically rely on biased modality-centric priors or enforce uniform retention policies, neglecting the heterogeneity across layers and the critical role of cross-modality alignment. To address these challenges, we propose OmniFit, a training-free framework that decouples interaction profiling from inference execution. OmniFit incorporates Layer-Adaptive Heterogeneity Profiling (LAHP) to dynamically allocate computational budgets based on layer-wise redundancy and modality preferences, preserving tokens according to the characteristics of each layer. Furthermore, we introduce Alignment-Rectified Token Selection (ARTS), a lightweight mechanism that efficiently identifies tokens semantically aligned with cross-modal cues. Extensive experiments on 3 model series across 10 benchmarks demonstrate that OmniFit establishes a new Pareto frontier, retaining 98% of model performance with only 20% token usage and achieves up to 2.31 end-to-end inference speedup and 2.5 VRAM saving, significantly outperforming state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 117b8fed-32d7-4d5d-a9e7-df7527eb3801Builds on19
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 451 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You et al.NeurIPS 2025 · 72 citations
Related papers
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language ModelsKeda Tao, Kele Shao, Bohan Yu, Weiqiang Wang et al.CVPR 2026 · 32 citations
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language ModelsYue Ding, Yiyan Ji, Jungang Li, Xuyang Liu et al.ICML 2026 · 22 citations
- Training-Free Multimodal Large Language Model OrchestrationTianyu Xie, Yuexiao Ma, Yuhang Wu, Wang Chen et al.ICML 2026 · 2 citations
- One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMsYongru Chen, Kai Zhang, Zeliang Zong, Yuchen Lu et al.CVPR 2026 · 1 citation
- SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language ModelsSeungil Lee, Gilha lee, Hyun KimICML 2026
