Lune

ACL2022Top-tier venue

Finding Structural Knowledge in Multimodal-BERT

Victor Milewski, Miryam de Lhoneux, Marie-Francine Moens

2022Year
12Citations
4Top-tier citations

Abstract

In this work, we investigate the knowledge learned in the embeddings of multimodal-BERT models. More specifically, we probe their capabilities of storing the grammatical structure of linguistic data and the structure learned over objects in visual data. To reach that goal, we first make the inherent structure of language and visuals explicit by a dependency parse of the sentences that describe the image and by the dependencies between the object regions in the image, respectively. We call this explicit visual structure the scene tree, that is based on the dependency tree of the language description. Extensive probing experiments show that the multimodal-BERT models do not encode these scene trees. Code available at https://github. com/VSJMilewski/multimodal-probes . Research Questions In this study, we aim to answer the following research questions.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 09a694c4-ed80-483d-b3b8-7c0aba8f1b49

Cited by top-tier papers4

Ask how each one uses it

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines