Video Scene Graph Generation from Single-Frame Weak Supervision
Siqi Chen, Jun Xiao, Long Chen
Abstract
Video scene graph generation (VidSGG) aims to generate a sequence of graph-structure representations for the given video. However, all existing VidSGG methods are fully-supervised, i.e., they need dense and costly manual annotations. In this paper, we propose the first weakly-supervised VidSGG task with only single-frame weak supervision: SF-VidSGG. By ``weakly-supervised", we mean that SF-VidSGG relaxes the training supervision from two different levels: 1) It only provides single-frame annotations instead of all-frame annotations. 2) The single-frame ground-truth annotation is still a weak image SGG annotation, i.e., an unlocalized scene graph. To solve this new task, we also propose a novel Pseudo Label Assignment based method, dubbed as PLA. PLA is a two-stage method, which generates pseudo visual relation annotations for the given video at the first stage, and then trains a fully-supervised VidSGG model with these pseudo labels. Specifically, PLA consists of three modules: an object PLA module, a predicate PLA module, and a future predicate prediction (FPP) module. Firstly, in the object PLA, we localize all objects for every frame. Then, in the predicate PLA, we design two different teachers to assign pseudo predicate labels. Lastly, in the FPP module, we fusion these two predicate pseudo labels by the regularity of relation transition in videos. Extensive ablations and results on the benchmark Action Genome have demonstrated the effectiveness of our PLA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1b2fb00f-fced-4ca2-97a9-5880ef465927Cited by top-tier papers7
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation DetectionKaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao et al.ICLR 2023 · 9 citations
- Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationWenqing Wang, Kaifeng Gao, Yawei Luo, Tao Jiang et al.ACM MM 2023 · 7 citations
- Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship DetectionShuo Yang, Yongqi Wang, Xiaofeng Ji, Xinxiao WuAAAI 2024 · 4 citations
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge TransferringZhu Xu, Ting Lei, Zhimin Li, Guan Wang et al.ICCV 2025 · 3 citations
- End-to-End Entity-Predicate Association Reasoning for Dynamic Scene Graph GenerationLiwei Wang, Yanduo Zhang, Tao Lu, Fang Liu et al.ICCV 2025 · 1 citation
Related papers
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 5 citations
- Weakly Supervised Video Scene Graph Generation via Natural Language SupervisionKibum Kim, Kanghoon Yoon, Yeonjun In, Jaehyeong Jeon et al.ICLR 2025
- A Simple Baseline for Weakly-Supervised Scene Graph GenerationJing Shi, Yiwu Zhong, Ning Xu, Yin Li et al.ICCV 2021 · 34 citations
- Prior Knowledge-driven Dynamic Scene Graph Generation with Causal InferenceJiale Lu, Lianggangxu Chen, Youqi Song, Shaohui Lin et al.ACM MM 2023 · 7 citations
- Semi-Supervised Clustering Framework for Fine-grained Scene Graph GenerationJiarui Yang, Chuan Wang, Jun Zhang, Shuyi Wu et al.AAAI 2025 · 2 citations
