AutoM3L: An Automated Multimodal Machine Learning Framework with Large Language Models
Daqin Luo, Chengjian Feng, Yuxuan Nong, Yiqing Shen
Abstract
Automated Machine Learning (AutoML) offers a promising approach to streamline the training of machine learning models. However, existing AutoML frameworks are often limited to unimodal scenarios and require extensive manual configuration. Recent advancements in Large Language Models (LLMs) have showcased their exceptional abilities in reasoning, interaction, and code generation, presenting an opportunity to develop a more automated and user-friendly framework. To this end, we introduce AutoM 3 L, an innovative Automated Multimodal Machine Learning framework that leverages LLMs as controllers to automatically construct multimodal training pipelines. AutoM 3 L comprehends data modalities and selects appropriate models based on user requirements, providing automation and interactivity. By eliminating the need for manual feature engineering and hyperparameter optimization, our framework simplifies user engagement and enables customization through directives, addressing the limitations of previous rule-based AutoML approaches. We evaluate the performance of AutoM 3 L on six diverse multimodal datasets spanning classification, regression, and retrieval tasks, as well as a comprehensive set of unimodal datasets. The results demonstrate that AutoM 3 L achieves competitive or superior performance compared to traditional rule-based AutoML methods. Furthermore, a user study highlights the user-friendliness and usability of our framework, compared to the rule-based AutoML methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab24edfa-e93a-4d89-8c02-9884d262870bCited by top-tier papers7
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng et al.NeurIPS 2025 · 20 citations
- AutoReproduce: Automatic AI Experiment Reproduction with Paper LineageXuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi et al.ACL 2026 · 18 citations
- RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous DrivingZhijian Huang, Chengjian Feng, Feng Yan, Baihui Xiao et al.ICCV 2025 · 6 citations
- AutoCT: Automating Interpretable Clinical Trial Prediction with LLM AgentsFengze Liu, Haoyu Wang, Joonhyuk Cho, Dan Roth et al.EMNLP 2025 · 1 citation
- RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and PredictionYufeng Zhong, Chengjian Feng, Feng Yan, Fanfan Liu et al.ICCV 2025 · 1 citation
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li et al.NeurIPS 2023 · 1,778 citations
- Deconfounded Multimodal Learning for Spatio-temporal Video GroundingJiawei Wang, Zhanchang Ma, Da Cao, Yuquan Le et al.ACM MM 2023 · 7 citations
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du et al.ICLR 2023
- MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkSirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng et al.ICLR 2024
Related papers
- AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoMLPatara Trirat, Wonyong Jeong, Sung Ju HwangICML 2025
- MLZero: A Multi-Agent System for End-to-end Machine Learning AutomationHaoyang Fang, Boran Han, Nick Erickson, Xiyuan Zhang et al.NeurIPS 2025 · 29 citations
- Mordal: Automated Pretrained Model Selection for Vision Language ModelsShiqi He, Insu Jang, Mosharaf ChowdhuryICLR 2026
- CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement LearningShun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie et al.AAAI 2026
- Towards Robust Multi-Modal Reasoning via Model SelectionXiangyan Liu, Rongxue Li, Wei Ji, Tao LinICLR 2024 · 9 citations
