Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
Xin Hu, Ke Qin, Wen Yin, Yuan-Fang Li, Ming Li, Tao He
Abstract
Scene Graph Generation (SGG) unifies object localization and visual relationship reasoning by predicting boxes and subject–predicate–object triples. Yet most pipelines treat SGG as a one-shot, deterministic classification instead of a genuine progressive, generative task. We propose FlowSG, which recasts SGG as continuous-time transport on a hybrid discrete–continuous state: starting from a noised graph, the model progressively grows an image-conditioned scene graph through constraint-aware refinements that jointly synthesize nodes (objects) and edges (predicates). Specifically, we first leverage a VQ-VAE to quantize a scene graph (e.g., the continuous visual features) into compact, predictable tokens; a graph Transformer then (i) predicts a conditional velocity field to transport continuous geometry (boxes) and (ii) updates discrete posteriors for categorical tokens (object features and predicate labels), coupling semantics and geometry via flow-conditioned message aggregation. Training combines flow-matching losses for geometry with a discrete-flow objective for tokens, yielding few-step inference and plug-and-play compatibility with standard detectors/segmenters. Extensive experiments on VG and PSG under closed- and open-vocabulary protocols show consistent gains in predicate R/mR and graph-level metrics, validating the mixed discrete–continuous generative formulation over one-shot classification baselines, e.g., an average improvement of about 3 points over the SOTA USG-Par.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ee6aeb4-3d4c-4169-b966-17b69f6c608eBuilds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- UniQ: Unified Decoder with Task-specific Queries for Efficient Scene Graph GenerationXinyao Liao, Wei Wei, Dangyang Chen, Yuanyuan FuACM MM 2024 · 2 citations
- Iterative Scene Graph GenerationSiddhesh Khandelwal, Leonid SigalNeurIPS 2022 · 47 citations
- Weakly Supervised Visual Semantic ParsingAlireza Zareian, Svebor Karaman, Shih-Fu ChangCVPR 2020
- Vision Relation Transformer for Unbiased Scene Graph GenerationGopika Sudhakaran, Devendra Singh Dhami, Kristian Kersting, Stefan RothICCV 2023 · 27 citations
- From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsRongjie Li, Songyang Zhang, Dahua Lin, Kai Chen et al.CVPR 2024
