On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
Geewook Kim, Minjoon Seo
摘要
Recent advancements in language and vision assistants have showcased impressive capabilities but suffer from a lack of transparency, limiting broader research and reproducibility. While open-source models handle general image tasks effectively, they face challenges with the high computational demands of complex visuallysituated text understanding. Such tasks often require increased token inputs and large vision modules to harness high-resolution information. Striking a balance between model size and data importance remains an open question. This study aims to redefine the design of visionlanguage models by identifying key components and creating efficient models with constrained inference costs. By strategically formulating datasets, optimizing vision modules, and enhancing supervision techniques, we achieve significant improvements in inference throughput while maintaining high performance. Extensive experiments across models ranging from 160M to 13B parameters offer insights into model optimization. We will fully opensource our codebase, models, and datasets at https://github.com/naver-ai/elva .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal ModelsGeewook Kim, Minjoon SeoAAAI 2026 · 被引用 1 次
- Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight MergingMinsik Choi, Geewook KimICML 2026 · 被引用 1 次
- How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?Seongyun Lee, Geewook Kim, Jiyeon Kim, Hyunji Lee 等ICLR 2025
- mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document UnderstandingAnwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye 等ACL 2025
它引用的顶会 Paper16
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactionsAdrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Yassine Ouali 等CVPR 2026
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras 等NeurIPS 2025 · 被引用 97 次
- Task-Aware Resolution Optimization for Visual Large Language ModelsWeiqing Luo, Zhen Tan, Yifan Li, Xinyu Zhao 等EMNLP 2025 · 被引用 1 次
- Inference Optimal VLMs Need Fewer Visual Tokens and More ParametersKevin Y. Li, Sachin Goyal, João D. Semedo, J. Zico KolterICLR 2025
