Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal Learning
Ziyue Wu, Junyu Gao, Changsheng Xu
Abstract
Video Scene Graph Generation (VidSGG), which aims to detect the relations between objects in a continuous spatio-temporal environment, has shown great potential in video understanding. Almost all prevailing VidSGG approaches are in a fully-supervised manner where expensive manual annotations are required. Therefore, we introduce a novel and challenging task named Weakly-supervised Video Scene Graph Generation (WS-VidSGG), in which a model is trained with only unlocalized scene graphs as supervisory information. Due to the imbalanced data distribution and the lack of fine-grained annotations, models learned in this setting is prone to be biased. Therefore, we propose an Unbiased Cross-Modal Learning (UCML) framework to address the WS-VidSGG task. Specifically, a cross-modal alignment module is firstly designed for allocating pseudo labels to unlabeled visual objects. We then extract unbiased knowledge from dataset statistics, and utilize prompt to make our model finely comprehend semantic concepts. The learned features that from the prompts and unbiased knowledge reinforced each other, resulting in discriminative textual representations. In order to better explore the relations between visual entities, we design a knowledge-guided attention graph to capture the cross-modal relations. Finally, the learned textual and visual features are integrated into a unified framework for relation prediction. Extensive ablation studies verify the effectiveness of our framework. Moreover, the comparison with state-of-the-art fully-supervised methods shows that our proposed framework also achieves comparable performance. Code https://github.com/ZiyueWu59/UCML is available.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Video Scene Graph Generation from Single-Frame Weak SupervisionSiqi Chen, Jun Xiao, Long ChenICLR 2023
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- Open-Vocabulary Video Scene Graph Generation via Union-aware Semantic AlignmentZiyue Wu, Junyu Gao, Changsheng XuACM MM 2024 · 8 citations
- Weakly-Supervised Video Object Grounding via Stable Context LearningWei Wang, Junyu Gao, Changsheng XuACM MM 2021 · 7 citations
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 5 citations
