Bridging the Sensory Gap: Visual Injection for Taxonomy Completion
Yuhang Niu, Hongyuan Xu, Ciyi Liu, Bofan Wei, Jiaqi Ye, Yanlong Wen, Xiaojie Yuan
Abstract
Taxonomy Completion aims to automatically integrate new concepts into existing hierarchies. However, existing text-only methods suffer from a "Sensory Gap": they struggle to differentiate ambiguous definitions (e.g., Latte vs. Cappuccino) and miss visual grouping signals. Consequently, they often misinterpret lexical overlaps as hierarchical dependencies, leading to erroneous structural predictions. To bridge this, we propose VITC, a framework leveraging Visual Injection for Taxonomy Completion. By mapping synthesized images into intrinsic pseudo-tokens, we enable the text encoder to perform holistic structural reasoning. To address injection challenges, we introduce Adaptive Residual Fusion, which decouples magnitude from selection to prevent visual signals from being drowned out, and the Multimodal Guided Adaptive Reweighting strategy, which leverages cross-modal consensus (Mutual Rescue and Complementary Mining) to filter noise and identify hard negatives. Experiments on three datasets demonstrate that VITC achieves state-of-the-art performance, delivering an average absolute gain of over 19%
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 052cba82-e900-4435-850e-38f83c712496Builds on22
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 999 citations
Related papers
- Towards Visual Taxonomy ExpansionTinghui Zhu, Jingping Liu, Jiaqing Liang, Haiyun Jiang et al.ACM MM 2023 · 3 citations
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 20 citations
- NativE: Multi-modal Knowledge Graph Completion in the WildYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.SIGIR 2024 · 39 citations
- Semantic-Augmented Image Clustering via Adaptive Multi-Modal CollaborationXiaohan Zhang, Chao Zhang, Deng Xu, Hong Yu et al.AAAI 2026
