Panoptic Video Scene Graph Generation
Jingkang Yang, Wenxuan Peng, Xiangtai Li, Zujin Guo, Liangyu Chen, Bo Li, Zheng Ma, Kaiyang Zhou, Wayne Zhang, Chen Change Loy, Ziwei Liu
2023Year
24Top-tier citations
Abstract
https://github.com/Jingkang50/OpenPVSG Frame 0000 Frame 0008 Frame 0042 Frame 0066 Frame 0104 Frame 0155 Frame 0194 Frame 0027 Frame 0360 Frame 0087 also provide a variety of baseline methods and share useful design practices for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 35 citations
- 4D Panoptic Scene Graph GenerationJingkang Yang, Jun Cen, Wenxuan Peng, Shuai Liu et al.NeurIPS 2023 · 33 citations
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi et al.CVPR 2024 · 14 citations
- CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial VideosTrong-Thuan Nguyen, Pha A. Nguyen, Xin Li, Jackson David Cothren et al.NeurIPS 2024 · 13 citations
Builds on21
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 398 citations
- SGTR: End-to-end Scene Graph Generation with TransformerRongjie Li, Songyang Zhang, Xuming HeCVPR 2022 · 108 citations
Related papers
- Relational Context Learning for Human-Object Interaction DetectionSanghyun Kim, Deunsol Jung, Minsu ChoCVPR 2023
- GaVS: 3D-Grounded Video Stabilization via Temporally-Consistent Local Reconstruction and RenderingZinuo You, Stamatios Georgoulis, Anpei Chen, Siyu Tang et al.SIGGRAPH 2025 · 3 citations
- FreeSim: Toward Free-viewpoint Camera Simulation in Driving ScenesLue Fan, Hao Zhang, Qitai Wang, Hongsheng Li et al.CVPR 2025
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern AnalysisMang Tik Chiu, Xingqian Xu, Yunchao Wei, Zilong Huang et al.CVPR 2020
- Human-Centric Multi-Exposure Fusion: Benchmark and Bi-level Cognition Distillation FrameworkJingjie Shang, Tengyu Ma, Heng Zhang, Jinyuan Liu et al.CVPR 2026
