PAVE: Patching and Adapting Video Large Language Models
Zhuoming Liu, Yiquan Li, Khoi Duc Nguyen, Yiwu Zhong, Yin Li
摘要
Pre-trained video large language models (Video LLMs) exhibit remarkable reasoning capabilities, yet adapting these models to new tasks involving additional modalities or data types (e.g., audio or 3D information) remains challenging. In this paper, we present PAVE, a flexible framework for adapting pre-trained Video LLMs to downstream tasks with side-channel signals, such as audio, 3D cues, or multi-view videos. PAVE introduces lightweight adapters, referred to as "patches," which add a small number of parameters and operations to a base model without modifying its architecture or pre-trained weights. In doing so, PAVE can effectively adapt the pre-trained base model to support diverse downstream tasks, including audio-visual question answering, 3D reasoning, multi-view video recognition, and high frame rate video understanding. Across these tasks, PAVE significantly enhances the performance of the base model, surpassing state-of-the-art task-specific models while incurring a minor cost of ∼0.1% additional FLOPs and parameters. Further, PAVE supports multitask learning and generalizes well across different Video LLMs. Our code is available at https://github. com/dragonlzm/PAVE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMsLidong Lu, Guo Chen, Zhu Wei, Zhiqi Li 等CVPR 2026 · 被引用 23 次
- V-LynX: Token Interface Alignment for Video+X LLMsJungin Park, Jiyoung Lee, Kwanghoon SohnICML 2026
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- LLMs Can Evolve Continually on Modality for X-Modal ReasoningJiazuo Yu, Haomiao Xiong, Lu Zhang, Haiwen Diao 等NeurIPS 2024 · 被引用 13 次
- TC-LLaVA: Rethinking the Transfer of LLava from Image to Video Understanding with Temporal ConsiderationsMingze Gao, Jingyu Liu, Mingda Li, Jiangtao Xie 等AAAI 2025 · 被引用 4 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- Parameter-Efficient Transfer Learning for Audio-Visual-Language TasksHongye Liu, Xianhai Xie, Yang Gao, Zhou YuACM MM 2023 · 被引用 2 次
- Toward Explainable Physical Audiovisual Commonsense ReasoningDaoming Zong, Chaoyue Ding, Kaitao ChenACM MM 2024
