Counterfactual Critic Multi-Agent Training for Scene Graph Generation
Long Chen, Hanwang Zhang, Jun Xiao, Xiangnan He, Shiliang Pu, Shih-Fu Chang
摘要
Scene graphs --- objects as nodes and visual relationships as edges --- describe the whereabouts and interactions of objects in an image for comprehensive scene understanding. To generate coherent scene graphs, almost all existing methods exploit the fruitful visual context by modeling message passing among objects. For example, person'' on bike'' can help to determine the relationship ``ride'', which in turn contributes to the confidence of the two objects. However, we argue that the visual context is not properly learned by using the prevailing cross-entropy based supervised learning paradigm, which is not sensitive to graph inconsistency: errors at the hub or non-hub nodes should not be penalized equally. To this end, we propose a Counterfactual critic Multi-Agent Training (CMAT) approach. CMAT is a multi-agent policy gradient method that frames objects into cooperative agents, and then directly maximizes a graph-level metric as the reward. In particular, to assign the reward properly to each agent, CMAT uses a counterfactual baseline that disentangles the agent-specific reward by fixing the predictions of other agents. Extensive validations on the challenging Visual Genome benchmark show that CMAT achieves a state-of-the-art performance by significant gains under various settings and metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- Model-Agnostic Counterfactual Reasoning for Eliminating Popularity Bias in Recommender SystemTianxin Wei, Fuli Feng, Jiawei Chen, Ziwei Wu 等KDD 2021 · 被引用 246 次
- Rethinking the Bottom-Up Framework for Query-Based Video LocalizationLong Chen, Chujie Lu, Siliang Tang, Jun Xiao 等AAAI 2020 · 被引用 182 次
- Reinforcement-Learning Based Portfolio Management with Augmented Asset Movement Prediction StatesYunan Ye, Hengzhi Pei, Boxin Wang, Pin-Yu Chen 等AAAI 2020 · 被引用 181 次
- Counterfactual Data-Augmented Sequential RecommendationZhenlei Wang, Jingsen Zhang, Hongteng Xu, Xu Chen 等SIGIR 2021 · 被引用 131 次
它引用的顶会 Paper1
相关 Paper
- Triangle-Reward Reinforcement Learning: A Visual-Linguistic Semantic Alignment for Image CaptioningWeizhi Nie, Jiesi Li, Ning Xu, An-An Liu 等ACM MM 2021 · 被引用 9 次
- Context-aware Scene Graph Generation with Seq2Seq TransformersYichao Lu, Himanshu Rai, Jason Chang, Boris Knyazev 等ICCV 2021 · 被引用 93 次
- Devil's on the Edges: Selective Quad Attention for Scene Graph GenerationDeunsol Jung, Sanghyun Kim, Won Hwa Kim, Minsu ChoCVPR 2023
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li 等ICCV 2021 · 被引用 34 次
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 被引用 47 次
