How Can Objects Help Video-Language Understanding?
Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo, Chen Sun
摘要
Do we still need to represent objects explicitly in multi-modal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves provide strong empirical performances for understanding tasks, despite missing fine-grained spatiotemporal information. To answer this question, we introduce ObjectMLLM, a framework capable of leveraging arbitrary computer vision algorithm to extract and integrate structured visual representation. Through extensive evaluations on six video question answering benchmarks, we confirm that explicit integration of object-centric representation remains necessary. Surprisingly, we observe that the simple approach of quantizing the continuous, structured object information and representing them as plain text performs the best, offering a data-efficient approach to integrate other visual perception modules into MLLM design. Our code and models are released at https://github.com/brown-palm/ObjectMLLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Object-Centric Concept-BottlenecksDavid Steinmann, Wolfgang Stammer, Antonia Wüst, Kristian KerstingNeurIPS 2025 · 被引用 12 次
- ViKey: Enhancing Temporal Understanding in Videos via Visual PromptingYeonkyung Lee, Dayun Ju, Youngmin Kim, Seil Kang 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
相关 Paper
- VideoOrion: Tokenizing Object Dynamics in VideosYicheng Feng, Yijiang Li, Wanpeng Zhang, Sipeng Zheng 等ICCV 2025
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision EncodersAli Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 等NeurIPS 2025 · 被引用 10 次
- Slot-VLM: Object-Event Slots for Video-Language ModelingJiaqi Xu, Cuiling Lan, Wenxuan Xie, Xuejin Chen 等NeurIPS 2024 · 被引用 13 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu 等ICML 2024 · 被引用 94 次
