Webly Supervised Knowledge Embedding Model for Visual Reasoning
Wenbo Zheng, Lan Yan, Chao Gou, Fei-Yue Wang
Abstract
Visual reasoning between visual image and natural language description is a long-standing challenge in computer vision. While recent approaches offer a great promise by compositionality or relational computing, most of them are oppressed by the challenge of training with datasets containing only a limited number of images with ground-truth texts. Besides, it is extremely time-consuming and difficult to build a larger dataset by annotating millions of images with text descriptions that may very likely lead to a biased model. Inspired by the majority success of webly supervised learning, we utilize readily-available web images with its noisy annotations for learning a robust representation. Our key idea is to presume on web images and corresponding tags along with fully annotated datasets in learning with knowledge embedding. We present a two-stage approach for the task that can augment knowledge through an effective embedding model with weakly supervised web data. This approach learns not only knowledge-based embeddings derived from key-value memory networks to make joint and full use of textual and visual information but also exploits the knowledge to improve the performance with knowledge-based representation learning for applying other general reasoning tasks. Experimental results on two benchmarks show that the proposed approach significantly improves performance compared with the state-ofthe-art methods and guarantees the robustness of our model against visual reasoning tasks and other reasoning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Open-Vocabulary Object Detection With an Open CorpusJiong Wang, Huiming Zhang, Haiwen Hong, Xuan Jin et al.ICCV 2023 · 22 citations
- Learning from the Past: Meta-Continual Learning with Knowledge Embedding for Jointly Sketch, Cartoon, and Caricature Face RecognitionWenbo Zheng, Lan Yan, Fei-Yue Wang, Chao GouACM MM 2020 · 19 citations
- lilGym: Natural Language Visual Reasoning with Reinforcement LearningAnne Wu, Kianté Brantley, Noriyuki Kojima, Yoav ArtziACL 2023
Builds on2
Related papers
- Beyond Embeddings: The Promise of Visual Table in Visual ReasoningYiwu Zhong, Zi-Yuan Hu, Michael R. Lyu, Liwei WangEMNLP 2024 · 1 citation
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
- Weakly-Supervised Image Hashing through Masked Visual-Semantic Graph-based ReasoningLu Jin, Zechao Li, Yonghua Pan, Jinhui TangACM MM 2020 · 30 citations
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou et al.ACM MM 2023 · 13 citations
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 5 citations
