Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine Comprehension
Huibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang, Yufan Li, Ning Jiang, Xin Wei, Zhenglu Yang
摘要
Procedural Multimodal Documents (PMDs) organize textual instructions and corresponding images step by step. Comprehending PMDs and inducing their representations for the downstream reasoning tasks is designated as Procedural MultiModal Machine Comprehension (M 3 C). In this study, we approach Procedural M 3 C at a fine-grained level (compared with existing explorations at a document or sentence level), that is, entity. With delicate consideration, we model entity both in its temporal and cross-modal relation and propose a novel Temporal-Modal Entity Graph (TMEG). Specifically, a heterogeneous graph structure is formulated to capture textual and visual entities and trace their temporal-modal evolution. In addition, a graph aggregation module is introduced to conduct graph encoding and reasoning. Comprehensive experiments across three Procedural M 3 C tasks are conducted on a traditional dataset RecipeQA and our new dataset CraftQA, which can better evaluate the generalization of TMEG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang 等ICCV 2019 · 被引用 441 次
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou 等ACL 2020 · 被引用 145 次
- A Dataset for Tracking Entities in Open Domain Procedural TextNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal 等EMNLP 2020 · 被引用 38 次
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 被引用 35 次
相关 Paper
- Reasoning over Entity-Action-Location Graph for Procedural Text UnderstandingHao Huang, Xiubo Geng, Jian Pei, Guodong Long 等ACL 2021
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu 等ACM MM 2020 · 被引用 20 次
- Procedural Text Understanding via Scene-Wise EvolutionJialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu 等AAAI 2022 · 被引用 2 次
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn 等EMNLP 2021
