Multi-Modal Few-Shot Temporal Action Segmentation
Zijia Lu, Ehsan Elhamifar
Abstract
Procedural videos are critical for learning new tasks. Temporal action segmentation (TAS), which classifies the action in every video frame, has become essential for understanding procedural videos. Existing TAS models, however, learn a fixed-set of tasks at training and unable to adapt to novel tasks at test time. Thus, we introduce the new problem of Multi-Modal Few-shot Temporal Action Segmentation (MMF-TAS) to learn open-set models that can generalize to novel procedural tasks with minimal visual/textual examples. We propose the first MMF-TAS framework, by designing a Prototype Graph Network (PGNet). In PGNet, a Prototype Building Block summarizes action information from support videos of the novel tasks via an Action Relation Graph, and encodes this information into action prototypes via a Dynamic Graph Transformer. Next, a Matching Block compares action prototypes with query videos to infer framewise action labels. To exploit the advantages of both visual and textual modalities, we compute separate action prototypes for each modality and combine the two modalities through prediction fusion to avoid overfitting on one modality. By extensive experiments on procedural datasets, we show our method successfully adapts to novel tasks during inference and significantly outperforms baselines. Our code is available at https://github.com/ZijiaLewisLu/ICCV2025-MMF-TAS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 4 citations
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 3 citations
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 3 citations
Builds on35
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang et al.ICCV 2023 · 113 citations
Related papers
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- Progress-Aware Online Action Segmentation for Egocentric Procedural Task VideosYuhan Shen, Ehsan ElhamifarCVPR 2024 · 14 citations
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu et al.CVPR 2022 · 73 citations
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- Adaptive FSS: A Novel Few-Shot Segmentation Framework via Prototype EnhancementJing Wang, Jiangyun Li, Chen Chen, Yisi Zhang et al.AAAI 2024 · 24 citations
