Attention-Based Context Aware Reasoning for Situation Recognition
Thilini Cooray, Ngai-Man Cheung, Wei Lu
摘要
Situation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action. Predicting semantic roles is very challenging: a vast variety of possibilities can be the match for a semantic role. Existing work has focused on dependency modelling architectures to solve this issue. Inspired by the success achieved by query-based visual reasoning (e.g., Visual Question Answering), we propose to address semantic role prediction as a query-based visual reasoning problem. However, existing query-based reasoning methods have not considered handling of inter-dependent queries which is a unique requirement of semantic role prediction in SR. Therefore, to the best of our knowledge, we propose the first set of methods to address inter-dependent queries in query-based visual reasoning. Extensive experiments demonstrate the effectiveness of our proposed method which achieves outstanding performance on Situation Recognition task. Furthermore, leveraging query inter-dependency, our methods improve upon a state-of-the-art method that answers queries separately.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue 等AAAI 2022 · 被引用 38 次
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang 等CVPR 2024 · 被引用 31 次
- Collaborative Transformers for Grounded Situation RecognitionJunhyeong Cho, Youngseok Yoon, Suha KwakCVPR 2022 · 被引用 23 次
- Graph-Wise Common Latent Factor Extraction for Unsupervised Graph Representation LearningThilini Cooray, Ngai-Man CheungAAAI 2022 · 被引用 5 次
相关 Paper
- Exploit Visual Dependency Relations for Semantic SegmentationMingyuan Liu, Dan Schonfeld, Wei TangCVPR 2021
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 被引用 38 次
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang 等CVPR 2022 · 被引用 89 次
- Explicit Knowledge Incorporation for Visual ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2021
- Mixture-Kernel Graph Attention Network for Situation RecognitionMohammed Suhail, Leonid SigalICCV 2019 · 被引用 32 次
