Open-World Attribute Mining for E-Commerce Products with Multimodal Self-Correction Instruction Tuning
Jiaqi Li, Yanming Li, Xiaoli Shen, Chuanyi Zhang, Guilin Qi, Sheng Bi
Abstract
In e-commerce, effective product Attribute Mining (AM) is essential for enhancing product features and aiding consumer decisions. However, current AM methods often focus on extracting attributes from unimodal text, underutilizing multimodal data. In this paper, we propose a novel framework called Multimodal Self-Correction Instruction Tuning (MSIT) to mine new potential attributes from images and texts with Multimodal Large Language Models (MLLMs). The tuning process involves two datasets: Attribute Generation Tuning Data (AGTD) and Chain-of-Thought Tuning Data (CTTD). AGTD is constructed utilizing incontext learning with a small set of seed attributes, aiding the MLLMs in accurately extracting attribute-value pairs from multimodal information. To introduce explicit reasoning and improve the extraction accuracy, we construct CTTD, which incorporates a structured 5-step reasoning process for self-correction. Finally, we employ a 3-stage inference process to filter out redundant attributes and sequentially validate each generated attribute. Comprehensive experimental results on two datasets show that MSIT outperforms state-of-the-art methods. We will release our code and data in the near future.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80e66501-e199-4277-9f47-65061cdddf6fCited by top-tier papers1
Ask how each one uses itBuilds on7
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng et al.NeurIPS 2021 · 1,026 citations
- OA-Mine: Open-World Attribute Mining for E-Commerce Products with Weak SupervisionXinyang Zhang, Chenwei Zhang, Xian Li, Xin Luna Dong et al.WWW 2022 · 36 citations
- TXtract: Taxonomy-Aware Knowledge Extraction for Thousands of Product CategoriesGiannis Karamanolakis, Jun Ma, Xin Luna DongACL 2020 · 3 citations
Related papers
- CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsXingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang et al.WWW 2025 · 15 citations
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah et al.KDD 2026 · 13 citations
- Prototype-Guided Multimodal Relation Extraction based on Entity AttributesZefan Zhang, Weiqi Zhang, Yanhui Li, Tian BaiAAAI 2025 · 8 citations
- Multimodal Joint Attribute Prediction and Value Extraction for E-commerce ProductTiangang Zhu, Yue Wang, Haoran Li, Youzheng Wu et al.EMNLP 2020 · 46 citations
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang et al.CVPR 2026 · 3 citations
