FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action Segmentation
Zijia Lu, Ehsan Elhamifar
Abstract
We study supervised action segmentation, whose goal is to predict framewise action labels of a video. To capture tem-poral dependencies over long horizons, prior works either improve framewise features with transformer or refine frame-wise predictions with learned action features. However, they are computationally costly and ignore that frame and action features contain complimentary information, which can be leveraged to enhance both features and improve temporal modeling. Therefore, we propose an efficient Frame-Action Cross-attention Temporal modeling (FACT) framework that performs temporal modeling withframe and action features in parallel and leverage this parallelism to achieve iterative bidirectional information transfer between the features and refine them. FACT network contains (i) aframe branch to learn frame-level information with convolutions and frame features, (ii) an action branch to learn action-level depen-dencies with transformers and action tokens and (iii) cross-attentions to allow communication between the two branches. We also propose a new matching loss to ensure each action to-ken uniquely encodes an action segment, thus better captures its semantics. Thanks to our architecture, we can also lever-age textual transcripts of videos to help action segmentation. We evaluate FACT on four video datasets (two egocentric and two third-person) for action segmentation with and without transcripts, showing that it significantly improves the state-of-the-art accuracy while enjoys lower computational cost (3 times faster) than existing transformer-based methods.<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>Code available at github.com/ZijiaLewisLu/CVPR2024-FACT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9a303e9-0289-417f-875d-bbacd7e72347Cited by top-tier papers19
- HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person ScenariosKunyu Peng, Junchao Huang, Xiangsheng Huang, Di Wen et al.NeurIPS 2025 · 12 citations
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 6 citations
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 3 citations
Builds on27
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- Dynamic DETR: End-to-End Object Detection with Dynamic AttentionXiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang et al.ICCV 2021 · 429 citations
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
Related papers
- Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary AlignmentAngchi Xu, Wei-Shi ZhengCVPR 2024 · 8 citations
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 35 citations
- Efficient Temporal Action Segmentation via Boundary-aware Query VotingPeiyao Wang, Yuewei Lin, Erik Blasch, Jie Wei et al.NeurIPS 2024 · 30 citations
- Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyZijia Lu, Ehsan ElhamifarCVPR 2022 · 22 citations
- How Much Temporal Long-Term Context is Needed for Action Segmentation?Emad Bahrami Rad, Gianpiero Francesca, Juergen GallICCV 2023 · 54 citations
