Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension
Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang
Abstract
Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M 3 C). In this study, we approach Procedural M 3 C at a fine-grained level (compared with existing explorations at a document or sentence level), that is, entity. With delicate consideration, we model entity both in its temporal and cross-modal relation and propose a novel Temporal-Modal Entity Graph (TMEG). Specifically, a heterogeneous graph structure is formulated to capture textual and visual entities and trace their temporal-modal evolution. In addition, a graph aggregation module is introduced to conduct graph encoding and reasoning. Comprehensive experiments across three Procedural M 3 C tasks are conducted on a traditional dataset RecipeQA and our new dataset CraftQA, which can better evaluate the generalization of TMEG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa8862de-a0f3-4d4f-bfb4-1897722f166aBuilds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
- A Dataset for Tracking Entities in Open Domain Procedural TextNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal et al.EMNLP 2020 · 38 citations
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
Related papers
- Reasoning over Entity-Action-Location Graph for Procedural Text UnderstandingHao Huang, Xiubo Geng, Jian Pei, Guodong Long et al.ACL 2021
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu et al.ACM MM 2020 · 20 citations
- Procedural Text Understanding via Scene-Wise EvolutionJialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu et al.AAAI 2022 · 2 citations
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 22 citations
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn et al.EMNLP 2021
