Markov Game Video Augmentation for Action Segmentation
Nicolas Aziere, Sinisa Todorovic
Abstract
This paper addresses data augmentation for action segmentation. Our key novelty is that we augment the original training videos in the deep feature space, not in the visual spatiotemporal domain as done by previous work. For augmentation, we modify original deep features of video frames such that the resulting embeddings fall closer to the class decision boundaries. Also, we edit action sequences of the original training videos (a.k.a. transcripts) by inserting, deleting, and replacing actions such that the resulting transcripts are close in edit distance to the ground-truth ones. For our data augmentation we resort to reinforcement learning, instead of more common supervised learning, since we do not have access to reliable oracles which would provide supervision about the optimal data modifications in the deep feature space. For modifying frame embeddings, we use a meta-model formulated as a Markov Game with multiple self-interested agents. Also, new transcripts are generated using a fast, parameter-free Monte Carlo tree search. Our experiments show that the proposed data augmentation of the Breakfast, GTEA, and 50Salads datasets leads to significant performance gains of several state of the art action segmenters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d66d03ca-b67f-4a3c-8ae0-3c6987672df4Cited by top-tier papers2
- FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationZijia Lu, Ehsan ElhamifarCVPR 2024 · 33 citations
- Action Sequence Augmentation for Action AnticipationYihui Qiu, Deepu RajanICLR 2025
Builds on8
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic MemoryLi Siyao, Weijiang Yu, Tianpei Gu, Chunze Lin et al.CVPR 2022 · 170 citations
- Refining Action Segmentation with Hierarchical Video RepresentationsHyemin Ahn, Dongheui LeeICCV 2021 · 74 citations
- Bridge-Prompt: Towards Ordinal Action Understanding in Instructional VideosMuheng Li, Lei Chen, Yueqi Duan, Zhilan Hu et al.CVPR 2022 · 70 citations
- Action Segmentation With Joint Self-Supervised Temporal Domain AdaptationMin-Hung Chen, Baopu Li, Yingze Bao, Ghassan AlRegib et al.CVPR 2020
Related papers
- Action Shuffle Alternating Learning for Unsupervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2021
- Fast Template Matching and Update for Video Object Tracking and SegmentationMingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang et al.CVPR 2020
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
- Data Augmentation for Instruction Following Policies via Trajectory SegmentationNiklas Höpner, Ilaria Tiddi, Herke van HoofAAAI 2025
- Coherent Temporal Synthesis for Incremental Action SegmentationGuodong Ding, Hans Golong, Angela YaoCVPR 2024
