OmniFit: Bridging Modalities via Layer-Adaptive Token Compression for Omnimodal Large Language Models
Zining Wang, Zhihang Yuan, Yingjie Zhai, Wenshuo Li, Han Shu, Ruihao Gong, Jinyang Guo, Xianglong Liu
摘要
Emerging Omni-modal Large Language Models (OmniLLMs) enable real-time interaction across video, audio, and text but suffer from prohibitive computational costs due to the quadratic complexity of processing continuous streaming inputs. Existing token compression strategies remain suboptimal as they typically rely on biased modality-centric priors or enforce uniform retention policies, neglecting the heterogeneity across layers and the critical role of cross-modality alignment. To address these challenges, we propose OmniFit, a training-free framework that decouples interaction profiling from inference execution. OmniFit incorporates Layer-Adaptive Heterogeneity Profiling (LAHP) to dynamically allocate computational budgets based on layer-wise redundancy and modality preferences, preserving tokens according to the characteristics of each layer. Furthermore, we introduce Alignment-Rectified Token Selection (ARTS), a lightweight mechanism that efficiently identifies tokens semantically aligned with cross-modal cues. Extensive experiments on 3 model series across 10 benchmarks demonstrate that OmniFit establishes a new Pareto frontier, retaining 98% of model performance with only 20% token usage and achieves up to 2.31 end-to-end inference speedup and 2.5 VRAM saving, significantly outperforming state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersHila Chefer, Shir Gur, Lior WolfICCV 2021 · 被引用 451 次
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang 等ICLR 2026 · 被引用 162 次
- HoliTom: Holistic Token Merging for Fast Video Large Language ModelsKele Shao, Keda Tao, Can Qin, Haoxuan You 等NeurIPS 2025 · 被引用 72 次
相关 Paper
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language ModelsKeda Tao, Kele Shao, Bohan Yu, Weiqiang Wang 等CVPR 2026 · 被引用 32 次
- OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language ModelsYue Ding, Yiyan Ji, Jungang Li, Xuyang Liu 等ICML 2026 · 被引用 22 次
- Training-Free Multimodal Large Language Model OrchestrationTianyu Xie, Yuexiao Ma, Yuhang Wu, Wang Chen 等ICML 2026 · 被引用 2 次
- One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMsYongru Chen, Kai Zhang, Zeliang Zong, Yuchen Lu 等CVPR 2026 · 被引用 1 次
- SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language ModelsSeungil Lee, Gilha lee, Hyun KimICML 2026
