MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action Anticipation
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca, Juergen Gall
Abstract
Long-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations. To model the uncertainty of future outcomes, stochastic models predict several potential future action sequences for the same observation. Recent work has further proposed to incorporate uncertainty modelling for observed frames by simultaneously predicting per-frame past and future actions in a unified manner. While such joint modelling of actions is beneficial, it requires long-range temporal capabilities to connect events across distant past and future time points. However, the previous work struggles to achieve such a long-range understanding due to its limited and/or sparse receptive field. To alleviate this issue, we propose a novel MANTA (MAmbafor ANTicipation) network. Our model enables effective long-term temporal modelling even for very long sequences while maintaining linear complexity in sequence length. We demonstrate that our approach achieves state-of-the-art results on three datasets—Breakfast, 50Salads, and Assembly 101—while also significantly improving computational and memory efficiency. Our code is available at https://github.com/olga-zats/DlFFMANTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd62fa1c-5cd7-4173-b0c8-e9cc49ff0171Cited by top-tier papers4
- RiverMamba: A State Space Model for Global River Discharge and Flood ForecastingMohamad Hakam Shams Eddin, Yikui Zhang, Stefan Kollet, Jürgen GallNeurIPS 2025 · 7 citations
- MS-Temba: Multi-Scale Temporal Mamba for Understanding Long Untrimmed VideosArkaprava Sinha, Monish Soundar Raj, Pu Wang, Ahmed Helmy et al.CVPR 2026 · 5 citations
- Action-Guided Attention for Video Action AnticipationTsung-Ming Tai, Sofia Casarin, Andrea Pilzer, Werner Nutt et al.ICLR 2026 · 5 citations
- Generative Video Compression with One-Dimensional Latent RepresentationZihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li et al.CVPR 2026 · 5 citations
Builds on23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
Related papers
- ActFusion: a Unified Diffusion Model for Action Segmentation and AnticipationDayoung Gong, Suha Kwak, Minsu ChoNeurIPS 2024 · 14 citations
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha et al.CVPR 2022 · 56 citations
- BiOMamba: Mamba-based Forward-Then-Backward Temporal Modeling for Online Action Detection and AnticipationSensen Wang, Yuehu Liu, Chi ZhangACM MM 2025
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- MixANT: Observation-Dependent Memory Propagation for Stochastic Dense Action AnticipationSyed Talal Wasim, Hamid Suleman, Olga Zatsarynna, Muzammal Naseer et al.ICCV 2025 · 1 citation
