4D Panoptic Scene Graph Generation
Jingkang Yang, Jun Cen, Wenxuan Peng, Shuai Liu, Fangzhou Hong, Xiangtai Li, Kaiyang Zhou, Qifeng Chen, Ziwei Liu
Abstract
We are living in a three-dimensional space while moving forward through a fourth dimension: time. To allow artificial intelligence to develop a comprehensive understanding of such a 4D environment, we introduce 4D Panoptic Scene Graph (PSG-4D), a new representation that bridges the raw visual data perceived in a dynamic 4D world and high-level visual understanding. Specifically, PSG-4D abstracts rich 4D sensory data into nodes, which represent entities with precise location and status information, and edges, which capture the temporal relations. To facilitate research in this new area, we build a richly annotated PSG-4D dataset consisting of 3K RGB-D videos with a total of 1M frames, each of which is labeled with 4D panoptic segmentation masks as well as fine-grained, dynamic scene graphs. To solve PSG-4D, we propose PSG4DFormer, a Transformer-based model that can predict panoptic segmentation masks, track masks along the time axis, and generate the corresponding scene graphs via a relation component. Extensive experiments on the new dataset show that our method can serve as a strong baseline for future research on PSG-4D. In the end, we provide a real-world application example to demonstrate how we can achieve dynamic scene understanding by integrating a large language model into our PSG-4D system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
- Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph GenerationThong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T. Nguyen et al.AAAI 2025 · 8 citations
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 5 citations
- PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous DrivingYining Pan, Shijie Li, Yuchen Wu, Xulei Yang et al.CVPR 2026 · 1 citation
- Unbiased Video Scene Graph Generation via Visual and Semantic Dual DebiasingYanjun Li, Zhaoyang Li, Honghui Chen, Lizhi XuCVPR 2025
Builds on25
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- Robust Multi-Modality Multi-Object TrackingWenwei Zhang, Hui Zhou, Shuyang Sun, Zhe Wang et al.ICCV 2019 · 221 citations
Related papers
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual SceneShengqiong Wu, Hao Fei, Jingkang Yang, Xiangtai Li et al.CVPR 2025
- RealGraph: A Multiview Dataset for 4D Real-world Context Graph GenerationHaozhe Lin, Zequn Chen, Jinzhi Zhang, Bing Bai et al.ICCV 2023 · 1 citation
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 48 citations
- X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferLinglin Jing, Ying Xue, Xu Yan, Chaoda Zheng et al.AAAI 2024 · 14 citations
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor SpacesChenyangguang Zhang, Alexandros Delitzas, Fangjinhua Wang, Ruida Zhang et al.CVPR 2025
