Webly Supervised Knowledge Embedding Model for Visual Reasoning
Wenbo Zheng, Lan Yan, Chao Gou, Fei-Yue Wang
摘要
Visual reasoning between visual image and natural language description is a long-standing challenge in computer vision. While recent approaches offer a great promise by compositionality or relational computing, most of them are oppressed by the challenge of training with datasets containing only a limited number of images with ground-truth texts. Besides, it is extremely time-consuming and difficult to build a larger dataset by annotating millions of images with text descriptions that may very likely lead to a biased model. Inspired by the majority success of webly supervised learning, we utilize readily-available web images with its noisy annotations for learning a robust representation. Our key idea is to presume on web images and corresponding tags along with fully annotated datasets in learning with knowledge embedding. We present a two-stage approach for the task that can augment knowledge through an effective embedding model with weakly supervised web data. This approach learns not only knowledge-based embeddings derived from key-value memory networks to make joint and full use of textual and visual information but also exploits the knowledge to improve the performance with knowledge-based representation learning for applying other general reasoning tasks. Experimental results on two benchmarks show that the proposed approach significantly improves performance compared with the state-ofthe-art methods and guarantees the robustness of our model against visual reasoning tasks and other reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Open-Vocabulary Object Detection With an Open CorpusJiong Wang, Huiming Zhang, Haiwen Hong, Xuan Jin 等ICCV 2023 · 被引用 22 次
- Learning from the Past: Meta-Continual Learning with Knowledge Embedding for Jointly Sketch, Cartoon, and Caricature Face RecognitionWenbo Zheng, Lan Yan, Fei-Yue Wang, Chao GouACM MM 2020 · 被引用 19 次
- lilGym: Natural Language Visual Reasoning with Reinforcement LearningAnne Wu, Kianté Brantley, Noriyuki Kojima, Yoav ArtziACL 2023
它引用的顶会 Paper2
相关 Paper
- Beyond Embeddings: The Promise of Visual Table in Visual ReasoningYiwu Zhong, Zi-Yuan Hu, Michael R. Lyu, Liwei WangEMNLP 2024 · 被引用 1 次
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- Weakly-Supervised Image Hashing through Masked Visual-Semantic Graph-based ReasoningLu Jin, Zechao Li, Yonghua Pan, Jinhui TangACM MM 2020 · 被引用 30 次
- Retrieval-based Knowledge Augmented Vision Language Pre-trainingJiahua Rao, Zifei Shan, Longpo Liu, Yao Zhou 等ACM MM 2023 · 被引用 13 次
- Web-Scale Visual Entity Recognition: An LLM-Driven Data ApproachMathilde Caron, Alireza Fathi, Cordelia Schmid, Ahmet IscenNeurIPS 2024 · 被引用 5 次
