Image Captioning with Multi-Context Synthetic Data
Feipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang, Xiaoyan Sun
Abstract
Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This potential can be harnessed to create synthetic image-text pairs for training captioning models. Synthetic data can improve cost and time efficiency in data collection, allow for customization to specific domains, bootstrap generalization capability for zero-shot performance, and circumvent privacy concerns associated with real-world data. However, existing methods struggle to attain satisfactory performance solely through synthetic data. We identify the issue as generated images from simple descriptions mostly capture a solitary perspective with limited context, failing to align with the intricate scenes prevalent in real-world imagery. To tackle this, we present an innovative pipeline that introduces multi-context data generation. Beginning with an initial text corpus, our approach employs a large language model to extract multiple sentences portraying the same scene from diverse viewpoints. These sentences are then condensed into a single sentence with multiple contexts. Subsequently, we generate intricate images using the condensed captions through diffusion models. Our model is exclusively trained on synthetic image-text pairs crafted through this process. The effectiveness of our pipeline is validated through experimental results in both the in-domain and cross-domain settings, where it achieves state-of-the-art performance on well-known datasets such as MSCOCO, Flickr30k, and NoCaps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2e1e629-6b6b-49eb-95c6-8ed9cf234550Cited by top-tier papers3
- IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot CaptioningSoeun Lee, Si-Woo Kim, Taewhan Kim, Dong-Jin KimEMNLP 2024 · 2 citations
- Engage for All: Making Ordinary Image Descriptions Appealing Again!Yuyan Chen, Yifan Jiang, Li Zhou, Jinghan Cao et al.ICCV 2025 · 1 citation
- SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image CaptioningSi-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee et al.ACM MM 2025
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- CapsFusion: Rethinking Image-Text Data at ScaleQiying Yu, Quan Sun, Xiaosong Zhang, Yufeng Cui et al.CVPR 2024 · 17 citations
- Diverse Image Captioning with Context-Object Split Latent SpacesShweta Mahajan, Stefan RothNeurIPS 2020 · 47 citations
- IT3D: Improved Text-to-3D Generation with Explicit View SynthesisYiwen Chen, Chi Zhang, Xiaofeng Yang, Zhongang Cai et al.AAAI 2024 · 80 citations
- JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation PromotionHaoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang et al.AAAI 2026
- MoS2: Mixture of Scale and Shift Experts for Text-Only Video CaptioningHeng Jia, Yunqiu Xu, Linchao Zhu, Guang Chen et al.ACM MM 2024 · 7 citations
