Predictive Feature Learning for Future Segmentation Prediction
Zihang Lin, Jiangxin Sun, Jianfang Hu, Qi-Zhi Yu, Jian-Huang Lai, Wei-Shi Zheng
Abstract
Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative (with rich details) and are always of high resolution/dimension. Hence, the complicated spatiotemporal variations of these features are difficult to predict, which motivates us to learn a more predictive representation. In this work, we develop a novel framework called Predictive Feature Autoencoder. In the proposed framework, we construct an autoencoder which serves as a bridge between the segmentation features and the predictor. In the latent feature learned by the autoencoder, global structures are enhanced and local details are suppressed so that it is more predictive. In order to reduce the risk of vanishing the suppressed details during recurrent feature prediction, we further introduce a reconstruction constraint in the prediction module. Extensive experiments show the effectiveness of the proposed approach and our method outperforms state-of-the-arts by a considerable margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fae905cb-2c18-442e-8e97-556fae173448Cited by top-tier papers5
- Real-time Object Detection for Streaming PerceptionJinrong Yang, Songtao Liu, Zeming Li, Xiaoping Li et al.CVPR 2022 · 61 citations
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- Temporal Continual Learning with Prior Compensation for Human Motion PredictionJianwei Tang, Jiangxin Sun, Xiaotong Lin, Lifang Zhang et al.NeurIPS 2023 · 6 citations
- Joint Forecasting of Panoptic Segmentations with Difference AttentionColin Graber, Cyril Jazra, Wenjie Luo, Liangyan Gui et al.CVPR 2022 · 3 citations
- Advancing Semantic Future Prediction through Multimodal Visual Sequence TransformersEfstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos KomodakisCVPR 2025
Builds on3
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Warp to the Future: Joint Forecasting of Features and Feature MotionJosip Saric, Marin Orsic, Tonci Antunovic, Sacha Vrazic et al.CVPR 2020
Related papers
- Seg-VAR: Image Segmentation with Visual Autoregressive ModelingRongkun Zheng, Lu Qi, Xi Chen, Yi Wang et al.NeurIPS 2025 · 3 citations
- Video Instance Segmentation Tracking With a Modified VAE ArchitectureChung-Ching Lin, Ying Hung, Rogério Feris, Linglin HeCVPR 2020
- Deep Hierarchical Video CompressionMing Lu, Zhihao Duan, Fengqing Zhu, Zhan MaAAAI 2024 · 19 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
