Understanding Multi-Task Activities from Single-Task Videos
Yuhan Shen, Ehsan Elhamifar
摘要
We introduce and develop a framework for Multi-Task Temporal Action Segmentation (MT-TAS), a novel paradigm that addresses the challenges of interleaved actions when performing multiple tasks simultaneously. Traditional action segmentation models, trained on single-task videos, struggle to handle task switches and complex scenes inherent in multi-task scenarios. To overcome these challenges, our MT-TAS approach synthesizes multi-task video data from single-task sources using our Multi-task Sequence Blending and Segment Boundary Learning modules. Additionally, we propose to dynamically isolate foreground and background elements within video frames, addressing the intricacies of object layouts in multi-task scenarios and enabling a new two-stage temporal action segmentation framework with Foreground-Aware Action Refinement. Also, we introduce the Multi-task Egocentric Kitchen Activities (MEKA) dataset, containing 12 hours of egocentric multi-task videos, to rigorously benchmark MT-TAS models. Extensive experiments demonstrate that our framework effectively bridges the gap between single-task training and multi-task testing, advancing temporal action segmentation with state-of-the-art performance in complex environments. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Multi-Modal Few-Shot Temporal Action SegmentationZijia Lu, Ehsan ElhamifarICCV 2025 · 被引用 6 次
- MOSCATO: Predicting Multiple Object State Change through ActionsParnian Zameni, Yuhan Shen, Ehsan ElhamifarICCV 2025 · 被引用 4 次
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 被引用 3 次
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 被引用 3 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural ActivitiesFadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He 等CVPR 2022 · 被引用 168 次
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang 等NeurIPS 2023 · 被引用 129 次
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang 等ICCV 2023 · 被引用 113 次
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 被引用 109 次
相关 Paper
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li 等AAAI 2025 · 被引用 3 次
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 被引用 74 次
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 被引用 26 次
- Tracking and Segmenting Anything in Any ModalityTianlu Zhang, Qiang Zhang, Guiguang Ding, Jungong HanAAAI 2026
