V2P: Vision-to-Prompt based Multi-Modal Product Summary Generation
Xuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao, Haiqing Chen, Liqiang Nie
摘要
Multi-modal Product Summary Generation is a new yet challenging task, which aims to generate a concise and readable summary for a product given its multi-modal content, e.g., its long text description and image. Although existing methods have achieved great success, they still suffer from three key limitations: 1) overlook the benefit of pre-training, 2) lack the representation-level supervision, and 3) ignore the diversity of the seller-generated data. To address these limitations, in this work, we propose a Vision-to-Prompt based multi-modal product summary generation framework, dubbed as V2P, where a Generative Pre-trained Language Model (GPLM) is adopted as the backbone. In particular, to maintain the original text capability of the GPLM and fully utilize the high-level concepts contained in the product image, we design V2P with two key components: vision-based prominent attribute prediction, and attribute prompt-guided summary generation. The first component works on obtaining the vital semantic attributes of the product from its image by the Swin Transformer, while the second component aims to generate the summary based on the product's long text description and the attribute prompts yielded by the first component with a GPLM. Towards comprehensive supervision over the second component, apart from the conventional output-level supervision, we introduce the representation-level regularization. Meanwhile, we design the data augmentation-based robustness regularization to handle the diverse inputs and improve the robustness of the second component. Extensive experiments on a large-scale Chinese dataset verify the superiority of our model over cutting-edge methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment AnalysisTeng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui 等ACM MM 2022 · 被引用 65 次
- Multi-Prompts Learning with Cross-Modal Alignment for Attribute-Based Person Re-identificationYajing Zhai, Yawen Zeng, Zhiyong Huang, Zheng Qin 等AAAI 2024 · 被引用 40 次
- Personalized Abstractive Opinion TaggingMengxue Zhao, Yang Yang, Miao Li, Jingang Wang 等SIGIR 2022 · 被引用 2 次
- MemSAM: Taming Segment Anything Model for Echocardiography Video SegmentationXiaolong Deng, Huisi Wu, Runhao Zeng, Jing QinCVPR 2024
相关 Paper
- CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained KnowledgeLinli Yao, Weijing Chen, Qin JinWWW 2023 · 被引用 11 次
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 被引用 16 次
- Adapting Generative Pretrained Language Model for Open-domain Multimodal Sentence SummarizationDengtian Lin, Liqiang Jing, Xuemeng Song, Meng Liu 等SIGIR 2023 · 被引用 15 次
- Concept-Guided Prompt Learning for Generalization in Vision-Language ModelsYi Zhang, Ce Zhang, Ke Yu, Yushun Tang 等AAAI 2024 · 被引用 37 次
