Panoptic Segmentation Forecasting
Colin Graber, Grace Tsai, Michael Firman, Gabriel J. Brostow, Alexander G. Schwing
Abstract
Our goal is to forecast the near future given a set of recent observations. We think this ability to forecast, i.e., to anticipate, is integral for the success of autonomous agents which need not only passively analyze an observation but also must react to it in real-time. Importantly, accurate forecasting hinges upon the chosen scene decomposition. We think that superior forecasting can be achieved by decomposing a dynamic scene into individual 'things' and background 'stuff'. Background 'stuff' largely moves because of camera motion, while foreground 'things' move because of both camera and individual object motion. Following this decomposition, we introduce panoptic segmentation forecasting. Panoptic segmentation forecasting opens up a middle-ground between existing extremes, which either forecast instance trajectories or predict the appearance of future image frames. To address this task we develop a twocomponent model: one component learns the dynamics of the background stuff by anticipating odometry, the other one anticipates the dynamics of detected things. We establish a leaderboard for this novel task, and validate a state-of-theart model that outperforms available baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72808a66-2837-4c5b-8f0a-cbce53f92a02Cited by top-tier papers3
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- Joint Forecasting of Panoptic Segmentations with Difference AttentionColin Graber, Cyril Jazra, Wenjie Luo, Liangyan Gui et al.CVPR 2022 · 3 citations
- Advancing Semantic Future Prediction through Multimodal Visual Sequence TransformersEfstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos KomodakisCVPR 2025
Builds on14
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- HarDNet: A Low Memory Traffic NetworkPing Chao, Chao-Yang Kao, Yu-Shan Ruan, Chien-Hsiang Huang et al.ICCV 2019 · 303 citations
- Self-Supervised Monocular Depth HintsJamie Watson, Michael Firman, Gabriel J. Brostow, Daniyar TurmukhambetovICCV 2019 · 287 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
Related papers
- Panoptic Neural Fields: A Semantic Object-Aware Neural Scene RepresentationAbhijit Kundu, Kyle Genova, Xiaoqi Yin, Alireza Fathi et al.CVPR 2022 · 204 citations
- Motion Forecasting in Continuous DrivingNan Song, Bozhou Zhang, Xiatian Zhu, Li ZhangNeurIPS 2024 · 33 citations
- Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersZhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu et al.CVPR 2022 · 145 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
- SDC-Depth: Semantic Divide-and-Conquer Network for Monocular Depth EstimationLijun Wang, Jianming Zhang, Oliver Wang, Zhe Lin et al.CVPR 2020
