Video Scene Graph Generation from Single-Frame Weak Supervision
Siqi Chen, Jun Xiao, Long Chen
摘要
Video scene graph generation (VidSGG) aims to generate a sequence of graph-structure representations for the given video. However, all existing VidSGG methods are fully-supervised, i.e., they need dense and costly manual annotations. In this paper, we propose the first weakly-supervised VidSGG task with only single-frame weak supervision: SF-VidSGG. By ``weakly-supervised", we mean that SF-VidSGG relaxes the training supervision from two different levels: 1) It only provides single-frame annotations instead of all-frame annotations. 2) The single-frame ground-truth annotation is still a weak image SGG annotation, i.e., an unlocalized scene graph. To solve this new task, we also propose a novel Pseudo Label Assignment based method, dubbed as PLA. PLA is a two-stage method, which generates pseudo visual relation annotations for the given video at the first stage, and then trains a fully-supervised VidSGG model with these pseudo labels. Specifically, PLA consists of three modules: an object PLA module, a predicate PLA module, and a future predicate prediction (FPP) module. Firstly, in the object PLA, we localize all objects for every frame. Then, in the predicate PLA, we design two different teachers to assign pseudo predicate labels. Lastly, in the FPP module, we fusion these two predicate pseudo labels by the regularity of relation transition in videos. Extensive ablations and results on the benchmark Action Genome have demonstrated the effectiveness of our PLA.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation DetectionKaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao 等ICLR 2023 · 被引用 9 次
- Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationWenqing Wang, Kaifeng Gao, Yawei Luo, Tao Jiang 等ACM MM 2023 · 被引用 7 次
- Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship DetectionShuo Yang, Yongqi Wang, Xiaofeng Ji, Xinxiao WuAAAI 2024 · 被引用 4 次
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge TransferringZhu Xu, Ting Lei, Zhimin Li, Guan Wang 等ICCV 2025 · 被引用 3 次
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu 等ICCV 2025 · 被引用 1 次
相关 Paper
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 被引用 5 次
- Weakly Supervised Video Scene Graph Generation via Natural Language SupervisionKibum Kim, Kanghoon Yoon, Yeonjun In, Jaehyeong Jeon 等ICLR 2025
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li 等ICCV 2021 · 被引用 34 次
- Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceJiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin 等ACM MM 2023 · 被引用 7 次
- Semi-Supervised Clustering Framework for Fine-grained Scene Graph GenerationJiarui Yang, Chuan Wang, Jun Zhang, Shuyi Wu 等AAAI 2025 · 被引用 2 次
