Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
Yang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu, Yuanxing Zhang, Zihao Wang, Weihong Lin, Jingyun Hua, Zekun Wang, Xinlong Chen, Bohan Zeng, Wentao Zhang
摘要
Long-context video understanding in Multimodal Large Language Models (MLLMs) faces a critical challenge: balancing computational efficiency with the retention of fine-grained spatio-temporal patterns. Existing approaches (e.g., sparse sampling, dense sampling with low resolution, and token compression) suffer from significant information loss in temporal dynamics, spatial details, or subtle interactions, particularly in videos with complex motion or varying resolutions. To address this, we propose Mavors, a novel framework that introduces Multi-granularity video representation for holistic long-video modeling. Specifically, Mavors directly encodes raw video content into latent representations through two core components: 1) an Intra-chunk Vision Encoder (IVE) that preserves high-resolution spatial features via 3D convolutions and Vision Transformers, and 2) an Inter-chunk Feature Aggregator (IFA) that establishes temporal coherence across chunks using transformer-based dependency modeling with chunk-level rotary position encodings. Moreover, the framework unifies image and video understanding by treating images as single-frame videos via sub-image decomposition. Experiments across diverse benchmarks demonstrate Mavors' superiority in maintaining both spatial fidelity and temporal continuity, significantly outperforming existing methods in tasks requiring fine-grained spatio-temporal reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System CollaborationHao Zhong, Muzhi Zhu, Zongze Du, Zheng Huang 等NeurIPS 2025 · 被引用 40 次
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang 等CVPR 2026 · 被引用 35 次
- VABench: A Comprehensive Benchmark for Audio-Video GenerationDaili Hua, Xizhi Wang, Bohan Zeng, Xinyi Huang 等CVPR 2026 · 被引用 25 次
- Debiasing Multimodal Large Language Models via Penalization of Language PriorsYifan Zhang, Yang Shi, Weichen Yu, Qingsong Wen 等ACM MM 2025 · 被引用 6 次
- Monet: Reasoning in Latent Visual Space Beyond Image and LanguageQixun Wang, Yang Shi, Yifei Wang, Yuanxing Zhang 等CVPR 2026
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video UnderstandingDaichi Yashima, Shuhei Kurita, Yusuke Oda, Komei SugiuraCVPR 2026 · 被引用 6 次
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu 等ICML 2025
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu 等ICML 2024 · 被引用 94 次
- Unifying Specialized Visual Encoders for Video Language ModelsJihoon Chung, Tyler Zhu, Max Gonzalez Saez-Diez, Juan Carlos Niebles 等ICML 2025
- Threading Keyframe with Narratives: MLLMs as Strong Long Video ComprehendersBo Fang, Yuxin Song, Haoyuan Sun, Qiangqiang Wu 等ICLR 2026 · 被引用 13 次
