DiffVsgg: Diffusion-Driven Online Video Scene Graph Generation
Mu Chen, Liulei Li, Wenguan Wang, Yi Yang
摘要
Top-leading solutions for Video Scene Graph Generation (VSGG) typically adopt an offline pipeline. Though demonstrating promising performance, they remain unable to handle real-time video streams and consume large GPU memory. Moreover, these approaches fall short in temporal reasoning, merely aggregating frame-level predictions over a temporal context. In response, we introduce DIFFVSGG, an online VSGG solution that frames this task as an iterative scene graph update problem. Drawing inspiration from Latent Diffusion Models (LDMs) which generate images via denoising a latent feature embedding, we unify the decoding of object classification, bounding box regression, and graph generation three tasks using one shared feature embedding. Then, given an embedding containing unified features of object pairs, we conduct a step-wise Denoising on it within LDMs, so as to deliver a clean embedding which clearly indicates the relationships between objects. This embedding then serves as the input to task-specific heads for object classification, scene graph generation, etc. DIF-FVSGG further facilitates continuous temporal reasoning, where predictions for subsequent frames leverage results of past frames as the conditional inputs of LDMs, to guide the reverse diffusion process for current frames. Extensive experiments on three setups of Action Genome demonstrate the superiority of DIFFVSGG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning Human-Object Interaction as GroupsJiajun Hong, Jianan Wei, Wenguan WangNeurIPS 2025 · 被引用 6 次
- What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token MergingInha Kang, Youngsun Lim, Seonho Lee, Jiho Choi 等ICLR 2026 · 被引用 1 次
- Reconciling Visual Perception and Generation in Diffusion ModelsLiulei Li, Yi Yang, Wenguan WangICLR 2026
- SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative AugmentationYiKai Li, Quhui Ke, Jinglin Liang, Zhiyuan Zhang 等ICML 2026
- BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship DetectionMelissa Schween, Mathis Kruse, Bodo RosenhahnCVPR 2026
它引用的顶会 Paper60
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 被引用 12 次
- Weakly Supervised Video Scene Graph Generation via Natural Language SupervisionKibum Kim, Kanghoon Yoon, Yeonjun In, Jaehyeong Jeon 等ICLR 2025
- Unifying Generation and Prediction on Graphs with Latent Graph DiffusionCai Zhou, Xiyuan Wang, Muhan ZhangNeurIPS 2024 · 被引用 37 次
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 被引用 57 次
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge TransferringZhu Xu, Ting Lei, Zhimin Li, Guan Wang 等ICCV 2025 · 被引用 3 次
