Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis
2025Year
2Top-tier citations
Abstract
5 IACM-Forth Inputs Predictions Oracle Figure 1. Our framework predicts future semantic segmentation and depth maps using a multimodal transformer architecture. Leveraging masked visual modeling and cross-modal fusion, it excels in future semantic prediction, achieving state-of-the-art results in both tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DINO-Foresight: Looking into the Future with DINOEfstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris, Nikos KomodakisNeurIPS 2025 · 52 citations
- Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionNils Morbitzer, Jonathan Evers, Artem Savkin, Thomas Stauner et al.ICML 2026
Builds on37
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
Related papers
- Multimodal Motion Prediction With Stacked TransformersYicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang et al.CVPR 2021
- Cross-view Transformers for real-time Map-view Semantic SegmentationBrady Zhou, Philipp KrähenbühlCVPR 2022 · 279 citations
- Masked-attention Mask Transformer for Universal Image SegmentationBowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov et al.CVPR 2022
- Joint Forecasting of Panoptic Segmentations with Difference AttentionColin Graber, Cyril Jazra, Wenjie Luo, Liangyan Gui et al.CVPR 2022 · 3 citations
- Matrix3D: Large Photogrammetry Model All-in-OneYuanxun Lu, Jingyang Zhang, Tian Fang, Jean-Daniel Nahmias et al.CVPR 2025
