Open-Vocabulary Video Scene Graph Generation via Union-aware Semantic Alignment
Ziyue Wu, Junyu Gao, Changsheng Xu
Abstract
Video Scene Graph Generation (VidSGG) plays a crucial role in various visual-language tasks by providing accessible structured visual relation knowledge. However, the requirement of annotating all categories of prevailing VidSGG methods limits their application in real-world scenarios. Despite the popular VLMs facilitating preliminary exploration of open-vocabulary VidSGG tasks, the correspondence between visual union regions and relation predicates is usually ignored. Therefore, we propose an Open-vocabulary VidSGG framework named Union-Aware Semantic Alignment Network (UASAN) to explore the alignment between visual union regions and relation predicate concepts in the same semantic space. Specifically, a visual refiner is designed to acquire open-vocabulary knowledge and the ability to bridge different modalities. To achieve better alignment, we first design a semantic-aware context encoder to achieve a comprehensive semantic interaction between object trajectories, visual union regions, and trajectory motion information to obtain semantic-aware union region representations. Then, a union-relation alignment decoder is utilized to generate the discriminative relation token for each union region for final relation prediction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 78051a8b-1224-458a-94de-1508e65edfa9Cited by top-tier papers2
- ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive ResponseXiaomeng ZHU, Fengming ZHU, Weijie Zhou, Ye Tian et al.ICML 2026
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneShengqiong Wu, Hao Fei, Jingkang Yang, Xiangtai Li et al.CVPR 2025
Related papers
- Weakly-supervised Video Scene Graph Generation via Unbiased Cross-modal LearningZiyue Wu, Junyu Gao, Changsheng XuACM MM 2023 · 5 citations
- Learning to Generate Language-Supervised and Open-Vocabulary Scene Graph Using Pre-Trained Visual-Semantic SpaceYong Zhang, Yingwei Pan, Ting Yao, Rui Huang et al.CVPR 2023
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
- Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph GenerationYiming Li, Sisi You, Bing-Kun BaoCVPR 2026
