ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
Dayoung Gong, Suha Kwak, Minsu Cho
摘要
Temporal action segmentation and long-term action anticipation are two popular vision tasks for the temporal analysis of actions in videos. Despite apparent relevance and potential complementarity, these two problems have been investigated as separate and distinct tasks. In this work, we tackle these two problems, action segmentation and action anticipation, jointly using a unified diffusion model dubbed ActFusion. The key idea to unification is to train the model to effectively handle both visible and invisible parts of the sequence in an integrated manner; the visible part is for temporal segmentation, and the invisible part is for future anticipation. To this end, we introduce a new anticipative masking strategy during training in which a late part of the video frames is masked as invisible, and learnable tokens replace these frames to learn to predict the invisible future. Experimental results demonstrate the bi-directional benefits between action segmentation and anticipation. ActFusion achieves the state-of-the-art performance across the standard benchmarks of 50 Salads, Breakfast, and GTEA, outperforming task-specific models in both of the two tasks with a single unified model through joint learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 被引用 6 次
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
- EmbryoDiff: A Conditional Diffusion Framework with Multi-Focal Feature Fusion for Fine-Grained Embryo Developmental Stage RecognitionYong Sun, Zhengjie Zhang, Junyu Shi, Zhiyuan Zhang 等AAAI 2026
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha 等CVPR 2022 · 被引用 56 次
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang 等ICCV 2023 · 被引用 113 次
- Predicting the Future: A Jointly Learnt Model for Action AnticipationHarshala Gammulle, Simon Denman, Sridha Sridharan, Clinton FookesICCV 2019 · 被引用 93 次
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action AnticipationOlga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca 等CVPR 2025
