Mixture-Kernel Graph Attention Network for Situation Recognition
Mohammed Suhail, Leonid Sigal
摘要
Understanding images beyond salient actions involves reasoning about scene context, objects, and the roles they play in the captured event. Situation recognition has recently been introduced as the task of jointly reasoning about the verbs (actions) and a set of semantic-role and entity (noun) pairs in the form of action frames. Labeling an image with an action frame requires an assignment of values (nouns) to the roles based on the observed image content. Among the inherent challenges are the rich conditional structured dependencies between the output role assignments and the overall semantic sparsity. In this paper, we propose a novel mixture-kernel attention graph neural network (GNN) architecture designed to address these challenges. Our GNN enables dynamic graph structure during training and inference, through the use of a graph attention mechanism, and context-aware interactions between role pairs. It also alleviates semantic sparsity by representing graph kernels using a convex combination of learned basis. We illustrate the efficacy of our model and design choices by conducting experiments on imSitu benchmark dataset, with accuracy improvements of up to 10% over state-of-the-art. 1 Semantic sparsity here refers to inability of a training dataset to span combinatorial number of possible action frame outputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SLAPS: Self-Supervision Improves Structure Learning for Graph Neural NetworksBahare Fatemi, Layla El Asri, Seyed Mehran KazemiNeurIPS 2021 · 被引用 220 次
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue 等AAAI 2022 · 被引用 38 次
- Collaborative Transformers for Grounded Situation RecognitionJunhyeong Cho, Youngseok Yoon, Suha KwakCVPR 2022 · 被引用 23 次
- ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision RepresentationYangyi Chen, Xingyao Wang, Manling Li, Derek Hoiem 等EMNLP 2023 · 被引用 2 次
相关 Paper
- Attention-Based Context Aware Reasoning for Situation RecognitionThilini Cooray, Ngai-Man Cheung, Wei LuCVPR 2020
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 被引用 57 次
- Devil's on the Edges: Selective Quad Attention for Scene Graph GenerationDeunsol Jung, Sanghyun Kim, Won Hwa Kim, Minsu ChoCVPR 2023
- InfoGCN: Representation Learning for Human Skeleton-based Action RecognitionHyung-Gun Chi, Myoung Hoon Ha, Seung-geun Chi, Sang Wan Lee 等CVPR 2022 · 被引用 383 次
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 被引用 19 次
