Hypergraph-based Zero-shot Multi-modal Product Attribute Value Extraction
Jiazhen Hu, Jiaying Gong, Hongda Shen, Hoda Eldardiry
摘要
It is essential for e-commerce platforms to provide accurate, complete, and timely product attribute values, in order to improve the search and recommendation experience for both customers and sellers. In the real-world scenario, it is difficult for these platforms to identify attribute values for the newly introduced products given no similar product history records for training or retrieval. Besides, how to jointly learn the product representation given various product information in multiple modalities, such as textual modality (e.g., product titles and descriptions) and visual modality (e.g., product images), is also a challenging task. To address these limitations, we propose a novel method for extracting multi-label product attribute-value pairs from multiple modalities in the zero-shot scenario, where labeled data is absent during training. Specifically, our method constructs heterogeneous hypergraphs, where product information from different modalities is represented by different types of nodes, and the text and image nodes are embedded and learned through CLIP encoders to effectively capture and integrate multi-modal product information. Then, the complex interrelations among these nodes are modeled through the hyperedges. By learning informative node representations, our method can accurately predict links between unseen product nodes and attribute-value nodes, enabling zero-shot attribute value extraction. We conduct extensive experiments and ablation studies on several categories of the public MAVE dataset and the results demonstrate that our proposed method significantly outperforms several state-of-the-art generative model baselines in multi-label, multi-modal product attribute value extraction in the zero-shot setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- LGMRec: Local and Global Graph Learning for Multimodal RecommendationZhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang 等AAAI 2024 · 被引用 164 次
相关 Paper
- Multi-Label Zero-Shot Product Attribute-Value ExtractionJiaying Gong, Hoda EldardiryWWW 2024 · 被引用 8 次
- Multimodal Joint Attribute Prediction and Value Extraction for E-commerce ProductTiangang Zhu, Yue Wang, Haoran Li, Youzheng Wu 等EMNLP 2020 · 被引用 46 次
- Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information ExtractionBaohang Zhou, Ying Zhang, Yu Zhao, Xuhui Sui 等WWW 2025 · 被引用 5 次
- Knowledge Perceived Multi-modal Pretraining in E-commerceYushan Zhu, Huaixiao Zhao, Wen Zhang, Ganqiang Ye 等ACM MM 2021 · 被引用 21 次
- Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationLikang Wu, Zhi Li, Hongke Zhao, Zhefeng Wang 等KDD 2023 · 被引用 4 次
