Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
Letian Zhang, Quan Cui, Bingchen Zhao, Cheng Yang
Abstract
The success of multi-modal large language models (MLLMs) has been largely attributed to the large-scale training data. However, the training data of many MLLMs is unavailable due to privacy concerns. The expensive and labor-intensive process of collecting multi-modal data further exacerbates the problem. Is it possible to synthesize multi-modal training data automatically without compromising diversity and quality? In this paper, we propose a new method, Oasis, to synthesize high-quality multi-modal data with only images. Oasis breaks through traditional methods by prompting only images to the MLLMs, thus extending the data diversity by a large margin. Our method features a delicate quality control method which ensures the data quality. We collected over 500k data and conducted incremental experiments on LLaVA-NeXT. Extensive experiments demonstrate that our method can significantly improve the performance of MLLMs. The image-based synthesis also allows us to focus on the specific-domain ability of MLLMs. Code and dataset are publicly available at https://github.com/Letian2003/MM_INF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Building a Foundational Guardrail for General Agentic Systems via Synthetic DataYue Huang, Hang Hua, Yujun Zhou, Pengcheng Jing et al.ICLR 2026 · 29 citations
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal PerceptionLai Wei, Liangbo He, jun lan, Lingzhong Dong et al.ICML 2026 · 27 citations
- ChemOrch: Empowering LLMs with Chemical Intelligence via Groundbreaking Synthetic InstructionsYue Huang, Zhengzhe Jiang, Xiaonan Luo, Kehan Guo et al.NeurIPS 2025 · 5 citations
- SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold StartKun Chen, Peng Shi, Haibo Qiu, Zhixiong Zeng et al.ICLR 2026 · 2 citations
- B-repLer: Language-guided Editing of CAD ModelsYilin Liu, Niladri Shekhar Dutt, Changjian Li, Niloy J. MitraSIGGRAPH 2026 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng et al.ICLR 2024 · 1,206 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
Related papers
- PANGEA: Projection-Based Augmentation with Non-Relevant General Data for Enhanced Domain Adaptation in LLMsSeungyoo Lee, Giung Nam, Moonseok Choi, Hyungi Lee et al.NeurIPS 2025
- CompCap: Improving Multimodal Large Language Models with Composite CaptionsXiaohui Chen, Satya Narayan Shukla, Mahmoud Azab, Aashu Singh et al.ICCV 2025 · 2 citations
- TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task TypesJiankang Chen, Tianke Zhang, Changyi Liu, Haojie Ding et al.ICLR 2025
- Img-Diff: Contrastive Data Synthesis for Multimodal Large Language ModelsQirui Jiao, Daoyuan Chen, Yilun Huang, Bolin Ding et al.CVPR 2025
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language ModelsHengzhuang Li, Xinsong Zhang, QIMING PENG, Bin Luo et al.CVPR 2026 · 2 citations
