Mordal: Automated Pretrained Model Selection for Vision Language Models
Shiqi He, Insu Jang, Mosharaf Chowdhury
摘要
Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, enabling them to perform multimodal tasks. Vision language models (VLMs) form the fastest growing category of multimodal models because of their many practical use cases, including in healthcare, robotics, and accessibility. Unfortunately, even though different VLMs in the literature demonstrate impressive visual capabilities in different benchmarks, they are handcrafted by human experts; there is no automated framework to create task-specific multimodal models.
We introduce Mordal, an automated multimodal model search framework that efficiently finds the best VLM for a user-defined task without manual intervention. Mordal achieves this both by reducing the number of candidates to consider during the search process and by minimizing the time required to evaluate each remaining candidate. Our evaluation shows that Mordal can find the best VLM for a given problem using -- lower GPU hours than grid search. We have also discovered that Mordal achieves about 69% higher weighted Kendall’s on average than the state-of-the-art model selection method across diverse tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language ModelsDaqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing ShenACM MM 2024 · 被引用 16 次
- Vision-Language Model Selection and Reuse for Downstream AdaptationHao-Zhe Tan, Zhi Zhou, Yufeng Li, Lan-Zhe GuoICML 2025
- Q-MoE: Connector for MLLMs with Text-Driven RoutingHanzi Wang, Jiamin Ren, Yifeng Ding, Lei Ren 等ACM MM 2024 · 被引用 1 次
- Automated Model Discovery via Multi-modal & Multi-step PipelineJungMok Lee, Nam Hyeon-Woo, Moon Ye-Bin, Junhyun Nam 等NeurIPS 2025 · 被引用 3 次
- Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition LearningWei Li, Hehe Fan, Yongkang Wong, Yi Yang 等ICML 2024 · 被引用 49 次
