Finding Structural Knowledge in Multimodal-BERT
Victor Milewski, Miryam de Lhoneux, Marie-Francine Moens
Abstract
In this work, we investigate the knowledge learned in the embeddings of multimodal-BERT models. More specifically, we probe their capabilities of storing the grammatical structure of linguistic data and the structure learned over objects in visual data. To reach that goal, we first make the inherent structure of language and visuals explicit by a dependency parse of the sentences that describe the image and by the dependencies between the object regions in the image, respectively. We call this explicit visual structure the scene tree, that is based on the dependency tree of the language description. Extensive probing experiments show that the multimodal-BERT models do not encode these scene trees. Code available at https://github. com/VSJMilewski/multimodal-probes . Research Questions In this study, we aim to answer the following research questions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09a694c4-ed80-483d-b3b8-7c0aba8f1b49Cited by top-tier papers4
- MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & TasksLetitia Parcalabescu, Anette FrankACL 2023 · 15 citations
- Measuring Progress in Fine-grained Vision-and-Language UnderstandingEmanuele Bugliarello, Laurent Sartran, Aishwarya Agrawal, Lisa Anne Hendricks et al.ACL 2023 · 9 citations
- Cross-modal Attention Congruence Regularization for Vision-Language Relation AlignmentRohan Pandey, Rulin Shao, Paul Pu Liang, Ruslan Salakhutdinov et al.ACL 2023 · 3 citations
- @ CREPE: Can Vision-Language Foundation Models Reason Compositionally?Zixian Ma, Jerry Hong, Mustafa Omer Gul, Mona Gandhi et al.CVPR 2023
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
- Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal TransformersStella Frank, Emanuele Bugliarello, Desmond ElliottEMNLP 2021 · 36 citations
- Visually Grounded Compound PCFGsYanpeng Zhao, Ivan TitovEMNLP 2020 · 35 citations
Related papers
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 7 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Unsupervised Vision-Language Parsing: Seamlessly Bridging Visual Scene Graphs with Language Structures via Dependency RelationshipsChao Lou, Wenjuan Han, Yuhuan Lin, Zilong ZhengCVPR 2022 · 9 citations
- Self-Supervised Relationship ProbingJiuxiang Gu, Jason Kuen, Shafiq R. Joty, Jianfei Cai et al.NeurIPS 2020 · 18 citations
- Do Neural Language Models Show Preferences for Syntactic Formalisms?Artur Kulmizev, Vinit Ravishankar, Mostafa Abdou, Joakim NivreACL 2020 · 1 citation
