Collaborative Transformers for Grounded Situation Recognition
Junhyeong Cho, Youngseok Yoon, Suha Kwak
摘要
Grounded situation recognition is the task of predicting the main activity, entities playing certain roles within the activity, and bounding-box groundings of the entities in the given image. To effectively deal with this challenging task, we introduce a novel approach where the two processes for activity classification and entity estimation are interactive and complementary. To implement this idea, we propose Collaborative Glance-Gaze TransFormer (CoFormer) that consists of two modules: Glance transformer for activity classification and Gaze transformer for entity estimation. Glance transformer predicts the main activity with the help of Gaze transformer that analyzes entities and their relations, while Gaze transformer estimates the grounded entities by focusing only on the entities relevant to the activity predicted by Glance transformer. Our CoFormer achieves the state of the art in all evaluation metrics on the SWiG dataset. Training code and model weights are available at https://github.com/jhcho99/CoFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura 等ACM MM 2022 · 被引用 40 次
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li 等ACM MM 2023 · 被引用 31 次
- Grounded Video Situation RecognitionZeeshan Khan, C. V. Jawahar, Makarand TapaswiNeurIPS 2022 · 被引用 19 次
- Video Event Extraction via Tracking Visual States of ArgumentsGuang Yang, Manling Li, Jiajie Zhang, Xudong Lin 等AAAI 2023 · 被引用 14 次
- Training Multimedia Event Extraction With Generated Images and CaptionsZilin Du, Yunxin Li, Xu Guo, Yidan Sun 等ACM MM 2023 · 被引用 7 次
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Spatially Conditioned Graphs for Detecting Human-Object InteractionsFrederic Z. Zhang, Dylan Campbell, Stephen GouldICCV 2021 · 被引用 170 次
- Paint Transformer: Feed Forward Neural Painting with Stroke PredictionSonghua Liu, Tianwei Lin, Dongliang He, Fu Li 等ICCV 2021 · 被引用 106 次
相关 Paper
- Rethinking the Two-Stage Framework for Grounded Situation RecognitionMeng Wei, Long Chen, Wei Ji, Xiaoyu Yue 等AAAI 2022 · 被引用 38 次
- Detector-Free Weakly Supervised Group Activity RecognitionDongkeun Kim, Jinsung Lee, Minsu Cho, Suha KwakCVPR 2022 · 被引用 62 次
- RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction DetectionJihwan Park, Chanhyeong Yang, Jinyoung Park, Taehoon Song 等CVPR 2026
- End-to-End Human-Gaze-Target Detection with TransformersDanyang Tu, Xiongkuo Min, Huiyu Duan, Guodong Guo 等CVPR 2022 · 被引用 69 次
- Object-aware Gaze Target DetectionFrancesco Tonini, Nicola Dall'Asen, Cigdem Beyan, Elisa RicciICCV 2023 · 被引用 38 次
