PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance
Shangkun Sun, Ruyang Liu, Haoran Tang, Yixiao Ge, Haibo Lu, Jiankun Yang, Chen Li
摘要
The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most existing video LLMs cannot handle hour-long videos, while methods custom for long videos tend to be ineffective for shorter videos and images. In this paper, we identify the key issue as the redundant content in videos. To address this, we propose a novel pooling strategy that simultaneously achieves token compression and instruction-aware visual feature aggregation. Our model is termed Prompt-guided Pooling LLaVA, or PPLLaVA for short. Specifically, PPLLaVA consists of three core components: the CLIP-based visual-prompt alignment that extracts visual information relevant to the user's instructions, the prompt-guided pooling that compresses the visual sequence to arbitrary scales using convolution-style pooling, and the clip context extension designed for lengthy prompt common in visual dialogue. Extensive experiments have validated the performance of our model. With superior throughput, PPLLaVA achieves better results on image benchmarks as a video LLM, while achieving state-of-the-art performance across various video benchmarks, excelling in tasks ranging from caption generation to multiple-choice questions, and handling video lengths from seconds to hours.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowRuyang Liu, Shangkun Sun, Haoran Tang, Wei Gao 等ICCV 2025 · 被引用 15 次
- Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data AssignmentHaoyang Li, Fangcheng Fu, Sheng Lin, Hao Ge 等SIGMOD 2026 · 被引用 7 次
- One Trajectory, One Token: Grounded Video Tokenization Via Panoptic Sub-Object TrajectoryChenhao Zheng, Jieyu Zhang, Mohammadreza Salehi, Ziqi Gao 等ICCV 2025 · 被引用 7 次
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video UnderstandingXiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng 等ICCV 2025 · 被引用 6 次
- Video Spatial Reasoning with Object-Centric 3D RolloutHaoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang 等AAAI 2026 · 被引用 3 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
相关 Paper
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu 等ICML 2025
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai 等CVPR 2026 · 被引用 7 次
- PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language ModelsChenyu Yang, Xuan Dong, Xizhou Zhu, Weijie Su 等CVPR 2025
- Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingJingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang 等NeurIPS 2025 · 被引用 25 次
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui 等EMNLP 2024 · 被引用 231 次
