Lune

CVPR2025Top-tier venue

Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers

Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis

2025Year
2Top-tier citations

Abstract

5 IACM-Forth Inputs Predictions Oracle Figure 1. Our framework predicts future semantic segmentation and depth maps using a multimodal transformer architecture. Leveraging masked visual modeling and cross-modal fusion, it excels in future semantic prediction, achieving state-of-the-art results in both tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers2

Ask how each one uses it

Builds on37

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines