Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, Shih-Fu Chang
Abstract
Scene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects. For example, person'' on bike'' can help to determine the relationship ``ride'', which in turn contributes to the confidence of the two objects. However, we argue that the visual context is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes should not be penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach. CMAT is a multi-agent policy gradient method that frames objects into cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the predictions of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art performance by significant gains under various settings and metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c74e465f-d7d9-476b-8486-459b2fc91855Cited by top-tier papers35
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu et al.KDD 2021 · 246 citations
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao et al.AAAI 2020 · 182 citations
- Reinforcement-Learning Based Portfolio Management with Augmented Asset Movement Prediction StatesYunan Ye, Hengzhi Pei, Boxin Wang, Pin-Yu Chen et al.AAAI 2020 · 181 citations
- Counterfactual Data-Augmented Sequential RecommendationZhenlei Wang, Jingsen Zhang, Hongteng Xu, Xu Chen et al.SIGIR 2021 · 131 citations
Builds on1
Related papers
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu et al.ACM MM 2021 · 9 citations
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev et al.ICCV 2021 · 93 citations
- Devil's on the Edges: Selective Quad Attention for Scene Graph GenerationDeunsol Jung, Sanghyun Kim, Won Hwa Kim, Minsu ChoCVPR 2023
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li et al.ICCV 2021 · 34 citations
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 47 citations
