Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation
Yufei Guo, Jing Ma, Yixuan Dong, Tianlu Zhang, Shijie Yang, Yanlong Zang, Weijie Ding, Pinghua Gong, Jungong Han
摘要
Multimodal item embeddings are crucial for e-commerce item-to-item (I2I) retrieval, yet real-world product images often contain promotional overlays and background clutter that inject spurious visual cues and degrade retrieval robustness. This issue is particularly pronounced in MLRM-style pipelines, where a frozen vision encoder is connected to an LLM through a lightweight connector that must selectively aggregate visual tokens. We propose Text-Guided Q-Former (TGQ-Former), a text-guided visual representation learning framework that leverages structured metadata as semantic guidance for visual token extraction while preserving complementary visual evidence. Concretely, TGQ-Former employs a hybrid-query connector to disentangle metadata-anchored and exploratory visual streams, and introduces a lightweight reliability-aware Dual-Gated Vector Modulation module to adaptively calibrate their contributions under noisy inputs. Experiments on large-scale, real-world e-commerce datasets with full-pool retrieval show that TGQ-Former consistently outperforms strong connector baselines and end-to-end MLLMs. On average, it improves Hit Rate@100 (H@100) by 3.03%, demonstrating the effectiveness of text-guided visual encoding for robust multimodal retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Q-MoE: Connector for MLLMs with Text-Driven RoutingHanzi Wang, Jiamin Ren, Yifeng Ding, Lei Ren 等ACM MM 2024 · 被引用 1 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
- Aid: Adapting Image2video Diffusion Models for Instruction-Guided Video PredictionZhen Xing, Qi Dai, Zejia Weng, Zuxuan Wu 等ICCV 2025 · 被引用 4 次
- Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular UnderstandingZihao Jing, QIUHAO Zeng, Ruiyi Fang, Yan Sun 等ICLR 2026 · 被引用 3 次
- Referencing Where to Focus: Improving Visual Grounding with Referential QueryYabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin 等NeurIPS 2024 · 被引用 9 次
