Attention-Based Context Aware Reasoning for Situation Recognition
Thilini Cooray, Ngai-Man Cheung, Wei Lu
Abstract
Situation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action. Predicting semantic roles is very challenging: a vast variety of possibilities can be the match for a semantic role. Existing work has focused on dependency modelling architectures to solve this issue. Inspired by the success achieved by query-based visual reasoning (e.g., Visual Question Answering), we propose to address semantic role prediction as a query-based visual reasoning problem. However, existing query-based reasoning methods have not considered handling of inter-dependent queries which is a unique requirement of semantic role prediction in SR. Therefore, to the best of our knowledge, we propose the first set of methods to address inter-dependent queries in query-based visual reasoning. Extensive experiments demonstrate the effectiveness of our proposed method which achieves outstanding performance on Situation Recognition task. Furthermore, leveraging query inter-dependency, our methods improve upon a state-of-the-art method that answers queries separately.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4825d994-cd1a-4168-8321-ab41a28fd89aCited by top-tier papers7
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura et al.ACM MM 2022 · 40 citations
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue et al.AAAI 2022 · 38 citations
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang et al.CVPR 2024 · 31 citations
- Collaborative Transformers for Grounded Situation RecognitionJunhyeong Cho, Youngseok Yoon, Suha KwakCVPR 2022 · 23 citations
- Graph-Wise Common Latent Factor Extraction for Unsupervised Graph Representation LearningThilini Cooray, Ngai-Man CheungAAAI 2022 · 5 citations
Related papers
- Exploit Visual Dependency Relations for Semantic SegmentationMingyuan Liu, Dan Schonfeld, Wei TangCVPR 2021
- Beyond OCR + VQA: Involving OCR into the Flow for Robust and Accurate TextVQAGangyan Zeng, Yuan Zhang, Yu Zhou, Xiaomeng YangACM MM 2021 · 38 citations
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang et al.CVPR 2022 · 89 citations
- Explicit Knowledge Incorporation for Visual ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2021
- Mixture-Kernel Graph Attention Network for Situation RecognitionMohammed Suhail, Leonid SigalICCV 2019 · 32 citations
