LLaVA-MS-PIT: Multi-Modal Schema-Guided Progressive Instruction Tuning for Multi-Modal Event Extraction
Hui Zhang, Po Hu, Wei Emma Zhang
Abstract
The proliferation of multi-modal data on the internet has intensified the need for structured event understanding across textual and visual modalities. However, existing multi-modal event extraction models suffer from three major limitations: the absence of explicit event schema guidance, coarse-grained multi-modal alignment strategies, and reliance on heterogeneous, misaligned multi-modal training datasets. To address these issues, we propose LLaVA-MS-PIT, a Multi-modal Schema-Guided Progressive Instruction Tuning Framework that explicitly injects structured multi-modal event schema knowledge into the model before event extraction. Specifically, we introduce the textual event schema to establish the model’s prior knowledge of event concepts and enhance its ability to reason about event structures, while the visual event schema is employed to bridge the representation gap between textual and visual modalities at the event level, enabling unified and semantically aligned event representations across modalities. Moreover, to alleviate data scarcity and modality misalignment inherent in current benchmarks, we construct imSitu-MEE, a high-quality multi-modal parallel dataset generated and annotated through schema-guided procedures. Extensive experiments demonstrate that LLaVA-MS-PIT achieves competitive performance on multi-modal event extraction benchmarks, underscoring the effectiveness and necessity of schema-guided progressive instruction tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3389e4b-dd77-4d40-986b-d7075eee520dBuilds on12
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Prompt for Extraction? PAIE: Prompting Argument Interaction for Event Argument ExtractionYubo Ma, Zehao Wang, Yixin Cao, Mukai Li et al.ACL 2022 · 182 citations
- CLIP-Event: Connecting Text and Images with Event StructuresManling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou et al.CVPR 2022 · 103 citations
- Cross-media Structured Common Space for Multimedia Event ExtractionManling Li, Alireza Zareian, Qi Zeng, Spencer Whitehead et al.ACL 2020 · 87 citations
- Is a Large Language Model a Good Annotator for Event Extraction?Ruirui Chen, Chengwei Qin, Weifeng Jiang, Dongkyu ChoiAAAI 2024 · 65 citations
Related papers
- INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction TuningWujian Peng, Lingchen Meng, Yitong Chen, Yiweng Xie et al.NeurIPS 2025 · 7 citations
- LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsFeng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang et al.ICLR 2025
- Multimedia Event Extraction with LLM Knowledge EditingJiaao Yu, Yijing Lin, Zhipeng Gao, Xuesong Qiu et al.EMNLP 2025
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind et al.CVPR 2025
- EventGPT: Event Stream Understanding with Multimodal Large Language ModelsShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang et al.CVPR 2025
