Joint Forecasting of Panoptic Segmentations with Difference Attention
Colin Graber, Cyril Jazra, Wenjie Luo, Liangyan Gui, Alexander G. Schwing
Abstract
Forecasting of a representation is important for safe and effective autonomy. For this, panoptic segmentations have been studied as a compelling representation in recent work. However, recent state-of-the-art on panoptic segmentation forecasting suffers from two issues: first, individual object instances are treated independently of each other; second, individual object instance forecasts are merged in a heuristic manner. To address both issues, we study a new panoptic segmentation forecasting model that jointly forecasts all object instances in a scene using a transformer model based on ‘difference attention.’ It further refines the predictions by taking depth estimates into account. We evaluate the proposed model on the Cityscapes and AIODrive datasets. We find difference attention to be particularly suitable for forecasting because the difference of quantities like locations enables a model to explicitly reason about velocities and acceleration. Because of this, we attain state-of-the-art on panoptic segmentation forecasting metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7d4d72c-da96-4fb8-956d-320ff953a37aCited by top-tier papers2
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- Advancing Semantic Future Prediction through Multimodal Visual Sequence TransformersEfstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos KomodakisCVPR 2025
Builds on19
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 658 citations
- Disentangling Propagation and Generation for Video PredictionHang Gao, Huazhe Xu, Qi-Zhi Cai, Ruth Wang et al.ICCV 2019 · 90 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
Related papers
- Panoptic Segmentation ForecastingColin Graber, Grace Tsai, Michael Firman, Gabriel J. Brostow et al.CVPR 2021
- Panoptic SegFormer: Delving Deeper into Panoptic Segmentation with TransformersZhiqi Li, Wenhai Wang, Enze Xie, Zhiding Yu et al.CVPR 2022 · 145 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
- Panoptic, Instance and Semantic Relations: A Relational Context Encoder to Enhance Panoptic SegmentationShubhankar Borse, Hyojin Park, Hong Cai, Debasmit Das et al.CVPR 2022 · 17 citations
- Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance LearningJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu et al.ICCV 2023 · 11 citations
