Finding Structural Knowledge in Multimodal-BERT
Victor Milewski, Miryam de Lhoneux, Marie-Francine Moens
摘要
In this work, we investigate the knowledge learned in the embeddings of multimodal-BERT models. More specifically, we probe their capabilities of storing the grammatical structure of linguistic data and the structure learned over objects in visual data. To reach that goal, we first make the inherent structure of language and visuals explicit by a dependency parse of the sentences that describe the image and by the dependencies between the object regions in the image, respectively. We call this explicit visual structure the scene tree, that is based on the dependency tree of the language description. Extensive probing experiments show that the multimodal-BERT models do not encode these scene trees. Code available at https://github. com/VSJMilewski/multimodal-probes . Research Questions In this study, we aim to answer the following research questions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & TasksLetitia Parcalabescu, Anette FrankACL 2023 · 被引用 15 次
- Measuring Progress in Fine-grained Vision-and-Language UnderstandingEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks 等ACL 2023 · 被引用 9 次
- Cross-modal Attention Congruence Regularization for Vision-Language Relation AlignmentRohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov 等ACL 2023 · 被引用 3 次
- @ CREPE: Can Vision-Language Foundation Models Reason Compositionally?Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi 等CVPR 2023
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal TransformersStella Frank, Emanuele Bugliarello, Desmond ElliottEMNLP 2021 · 被引用 36 次
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 被引用 35 次
相关 Paper
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 被引用 7 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 被引用 9 次
- Self-Supervised Relationship ProbingJiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai 等NeurIPS 2020 · 被引用 18 次
- Do Neural Language Models Show Preferences for Syntactic Formalisms?Artur Kulmizev, Vinit Ravishankar, Mostafa Abdou, Joakim NivreACL 2020 · 被引用 1 次
