Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense Discovery
Feihong Lu, Weiqi Wang, Yangyifei Luo, Ziqin Zhu, Qingyun Sun, Baixuan Xu, Haochen Shi, Shiqi Gao, Qian Li, Yangqiu Song, Jianxin Li
摘要
Social media has become a ubiquitous tool for connecting with others, staying updated with news, expressing opinions, and finding entertainment. However, understanding the intention behind social media posts remains challenging due to the implicit and commonsense nature of these intentions, the need for cross-modality understanding of both text and images, and the presence of noisy information such as hashtags, misspelled words, and complicated abbreviations. To address these challenges, we present Mi o , a Multimodal Intention Kowledge DistillatiOn framework that collaboratively leverages a Large Language Model (LLM) and a Multimodal Large Language Model (MLLM) to uncover users' intentions. Specifically, we use an MLLM to interpret the image and an LLM to extract key information from the text and finally instruct the LLM again to generate intentions. By applying Mi o to publicly available social media datasets, we construct an intention knowledge base featuring 1,372K intentions rooted in 137,287 posts. We conduct a two-stage annotation to verify the quality of the generated knowledge and benchmark the performance of widely used LLMs for intention generation. We further apply Mi o to a sarcasm detection dataset and distill a student model to demonstrate the downstream benefits of applying intention knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag 等ACL 2025 · 被引用 15 次
- CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi 等ACL 2024 · 被引用 10 次
- MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase UnderstandingBaixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding 等EMNLP 2024 · 被引用 4 次
它引用的顶会 Paper13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Reasoning with Multimodal Sarcastic Tweets via Modeling Cross-Modality Contrast and Semantic AssociationNan Xu, Zhixiong Zeng, Wenji MaoACL 2020 · 被引用 153 次
- Data-efficient Fine-tuning for LLM-based RecommendationXinyu Lin, Wenjie Wang, Yongqi Li, Shuo Yang 等SIGIR 2024 · 被引用 152 次
- Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional NetworkBin Liang, Chenwei Lou, Xiang Li, Min Yang 等ACL 2022 · 被引用 151 次
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui 等ACM MM 2021 · 被引用 128 次
相关 Paper
- CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsZixin Chen, Hongzhan Lin, Ziyang Luo, Mingfei Cheng 等ACL 2024 · 被引用 10 次
- Mutual-Enhanced Incongruity Learning Network for Multi-Modal Sarcasm DetectionYang Qiao, Liqiang Jing, Xuemeng Song, Xiaolin Chen 等AAAI 2023 · 被引用 84 次
- S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionZhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun 等AAAI 2026
- LLM vs Small Model? Large Language Model Based Text Augmentation Enhanced Personality Detection ModelLinmei Hu, Hongyu He, Duokang Wang, Ziwang Zhao 等AAAI 2024 · 被引用 49 次
- Nice Perfume. How Long Did You Marinate in It? Multimodal Sarcasm ExplanationPoorav Desai, Tanmoy Chakraborty, Md. Shad AkhtarAAAI 2022 · 被引用 49 次
