Rethinking the Two-Stage Framework for Grounded Situation Recognition
Meng Wei, Long Chen, Wei Ji, Xiaoyu Yue, Tat-Seng Chua
摘要
Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g., buying) and detecting all corresponding semantic roles (e.g., agent and goods), is an essential step towards "human-like" event understanding. Since each verb is associated with a specific set of semantic roles, all existing GSR methods resort to a two-stage framework: predicting the verb in the first stage and detecting the semantic roles in the second stage. However, there are obvious drawbacks in both stages: 1) The widely-used cross-entropy (XE) loss for object recognition is insufficient in verb classification due to the large intraclass variation and high inter-class similarity among daily activities. 2) All semantic roles are detected in an autoregressive manner, which fails to model the complex semantic relations between different roles. To this end, we propose a novel SituFormer for GSR which consists of a Coarse-to-Fine Verb Model (CFVM) and a Transformer-based Noun Model (TNM). CFVM is a two-step verb prediction model: a coarse-grained model trained with XE loss first proposes a set of verb candidates, and then a fine-grained model trained with triplet loss re-ranks these candidates with enhanced verb features (not only separable but also discriminative). TNM is a transformer-based semantic role detection model, which detects all roles parallelly. Owing to the global relation modeling ability and flexibility of the transformer decoder, TNM can fully explore the statistical dependency of the roles. Extensive validations on the challenging SWiG benchmark show that SituFormer achieves a new state-of-the-art performance with significant gains under various metrics. Code is available at https://github.com/kellyiss/SituFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Panoptic Scene Graph Generation with Semantics-Prototype LearningLi Li, Wei Ji, Yiming Wu, Mengze Li 等AAAI 2024 · 被引用 63 次
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 被引用 19 次
- Video Event Extraction via Tracking Visual States of ArgumentsGuang Yang, Manling Li, Jiajie Zhang, Xudong Lin 等AAAI 2023 · 被引用 14 次
它引用的顶会 Paper9
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He 等ICCV 2019 · 被引用 165 次
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li 等AAAI 2022 · 被引用 145 次
- Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression GroundingLong Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang 等AAAI 2021 · 被引用 118 次
- Mixture-Kernel Graph Attention Network for Situation RecognitionMohammed Suhail, Leonid SigalICCV 2019 · 被引用 32 次
- HOSE-Net: Higher Order Structure Embedded Network for Scene Graph GenerationMeng Wei, Chun Yuan, Xiaoyu Yue, Kuo ZhongACM MM 2020 · 被引用 20 次
相关 Paper
- Collaborative Transformers for Grounded Situation RecognitionJunhyeong Cho, Youngseok Yoon, Suha KwakCVPR 2022 · 被引用 23 次
- Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language ExplainerJiaming Lei, Lin Li, Chunping Wang, Jun Xiao 等ACM MM 2024
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang 等EMNLP 2021 · 被引用 63 次
- RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction DetectionJihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song 等CVPR 2026
- Attention-Based Context Aware Reasoning for Situation RecognitionThilini Cooray, Ngai-Man Cheung, Wei LuCVPR 2020
