AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language Models
Daqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing Shen
摘要
Automated Machine Learning (AutoML) offers a promising approach to streamline the training of machine learning models. However, existing AutoML frameworks are often limited to unimodal scenarios and require extensive manual configuration. Recent advancements in Large Language Models (LLMs) have showcased their exceptional abilities in reasoning, interaction, and code generation, presenting an opportunity to develop a more automated and user-friendly framework. To this end, we introduce AutoM 3 L, an innovative Automated Multimodal Machine Learning framework that leverages LLMs as controllers to automatically construct multimodal training pipelines. AutoM 3 L comprehends data modalities and selects appropriate models based on user requirements, providing automation and interactivity. By eliminating the need for manual feature engineering and hyperparameter optimization, our framework simplifies user engagement and enables customization through directives, addressing the limitations of previous rule-based AutoML approaches. We evaluate the performance of AutoM 3 L on six diverse multimodal datasets spanning classification, regression, and retrieval tasks, as well as a comprehensive set of unimodal datasets. The results demonstrate that AutoM 3 L achieves competitive or superior performance compared to traditional rule-based AutoML methods. Furthermore, a user study highlights the user-friendliness and usability of our framework, compared to the rule-based AutoML methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng 等NeurIPS 2025 · 被引用 20 次
- AutoReproduce: Automatic AI Experiment Reproduction with Paper LineageXuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi 等ACL 2026 · 被引用 18 次
- RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous DrivingZhijian Huang, Chengjian Feng, Feng Yan, Baihui Xiao 等ICCV 2025 · 被引用 6 次
- AutoCT: Automating Interpretable Clinical Trial Prediction with LLM AgentsFengze Liu, Haoyu Wang, Joonhyuk Cho, Dan Roth 等EMNLP 2025 · 被引用 1 次
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Deconfounded Multimodal Learning for Spatio-temporal Video GroundingJiawei Wang, Zhanchang Ma, Da Cao, Yuquan Le 等ACM MM 2023 · 被引用 7 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
- MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkSirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng 等ICLR 2024
相关 Paper
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
- MLZero: A Multi-Agent System for End-to-end Machine Learning AutomationHaoyang Fang, Boran Han, Nick Erickson, Xiyuan Zhang 等NeurIPS 2025 · 被引用 29 次
- Mordal: Automated Pretrained Model Selection for Vision Language ModelsShiqi He, Insu Jang, Mosharaf ChowdhuryICLR 2026
- CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement LearningShun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie 等AAAI 2026
- Towards Robust Multi-Modal Reasoning via Model SelectionXiangyan Liu, Rongxue Li, Wei Ji, Tao LinICLR 2024 · 被引用 9 次
