mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, Jingren Zhou
2025Year
150Top-tier citations
Abstract
At the beginning of the movie, what does the policemen wear on their faces At the beginning of the movie, the policemen wear masks on their faces In the post-production segment of the film, what color is the car lifted by the bulldozer? The car lifted by the bulldozer is red in color.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fca95255-a10f-47ad-a8a4-2beb858735e3Cited by top-tier papers150
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMsJack Hong, Shilin Yan, Jiayin Cai, Xiaolong Jiang et al.ICLR 2026 · 162 citations
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-TuningQi (Cheems) Wang, Yanrui Yu, Ye Yuan, Rui Mao et al.NeurIPS 2025 · 103 citations
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingXiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li et al.NeurIPS 2025 · 95 citations
Builds on25
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Learning Multi-Object Tracking and Segmentation From Automatic AnnotationsLorenzo Porzi, Markus Hofinger, Idoia Ruiz, Joan Serrat et al.CVPR 2020
- Masked Face Recognition with Latent Part DetectionFeifei Ding, Peixi Peng, Yangru Huang, Mengyue Geng et al.ACM MM 2020 · 103 citations
- 3DEnhancer: Consistent Multi-View Diffusion for 3D EnhancementYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan et al.CVPR 2025
- Masked Face Recognition with Generative-to-Discriminative RepresentationsShiming Ge, Weijia Guo, Chenyu Li, Junzheng Zhang et al.ICML 2024 · 3 citations
- Learning Complete 3D Morphable Face Models From Images and VideosMallikarjun B. R., Ayush Tewari, Hans-Peter Seidel, Mohamed A. Elgharib et al.CVPR 2021
