REFINESUMM: Self-Refining MLLM for Generating a Multimodal Summarization Dataset
Vaidehi Patil, Leonardo F. R. Ribeiro, Mengwen Liu, Mohit Bansal, Markus Dreyer
Abstract
Multimodal Large Language Models (MLLMs) excel at synthesizing key information from diverse sources. However, generating accurate and faithful multimodal summaries is challenging, primarily due to the lack of appropriate multimodal datasets for fine-tuning that meaningfully integrate textual and visual modalities. To address this gap, we present a new dataset specifically designed for image-text multimodal summarization, harnessing the capabilities of state-of-the-art MLLMs. We generate summaries from Wikipedia sections and corresponding images and evaluate them across text-based, visual and multimodal dimensions, employing reference-free metrics. To refine the dataset, we: (1) filter the MLLM-generated summaries by training a critic model on human annotations and using its predictions to remove low-quality summaries; (2) fine-tune the MLLM with the filtered high-quality summaries; (3) use the fine-tuned model in turn to regenerate the summaries. This self-refinement process notably improves summary quality, as measured by human judgments and automatic multimodal metrics, resulting in a valuable dataset for multimodal summarization research. 1 * Work done as an intern at Amazon AGI. 1 The dataset is publicly available at https://github. com/amazon-science/refinesumm . The Italian wall lizard or ruin lizard (Podarcis siculus, from the Greek meaning agile and feet) is a species of lizard in the family Lacertidae. P. siculus is native to Bosnia and Herzegovina,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0b2e91a-66e8-48ea-8768-5812b1218f6aCited by top-tier papers4
- COSMMIC: Comment-Sensitive Multimodal Multilingual Indian Corpus for Summarization and Headline GenerationRaghvendra Kumar, Mohammed Salman S. A, Aryan Sahu, Tridib Nandi et al.ACL 2025
- What Is That Talk About? A Video-to-Text Summarization Dataset for Scientific PresentationsDongqi Liu, Chenxi Whitehouse, Xi Yu, Louis Mahon et al.ACL 2025
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang et al.ACL 2025
- Bootstrapping Language-Guided Navigation Learning with Self-Refining Data FlywheelZun Wang, Jialu Li, Yicong Hong, Songze Li et al.ICLR 2025
Builds on17
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li et al.ICLR 2026 · 4 citations
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 5 citations
- CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and GenerationWei Chen, Lin Li, Yongqi Yang, Bin Wen et al.CVPR 2025
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical ImagesGuimeng Liu, Tianze Yu, Somayeh Ebrahimkhani, Lin Zhi Zheng Shawn et al.ICLR 2026 · 3 citations
