Hierarchical Action Learning for Weakly-Supervised Action Segmentation
Junxian Huang, Ruichu Cai, Juntao Fang, Hao Zhu, Boyan Xu, Weilin Chen, Zijian Li, Shenghua Gao
摘要
Humans perceive actions through key transitions that structure actions across multiple abstraction levels, whereas machines, relying on visual features, tend to over-segment. This highlights the difficulty of enabling hierarchical reasoning in video understanding. Interestingly, we observe that lower-level visual and high-level action latent variables evolve at different rates, with low-level visual variables changing rapidly, while high-level action variables evolve more slowly, making them easier to identify. Building on this insight, we propose the Hierarchical Action Learning (HAL) model for weakly-supervised action segmentation. Our approach introduces a hierarchical causal data generation process, where high-level latent action govern the dynamics of low-level visual features. To model these varying timescales effectively, we introduce deterministic processes to align these latent variables over time. The HAL model employs a hierarchical pyramid transformer to capture both visual features and latent variables, and a sparse transition constraint is applied to enforce the slower dynamics of high-level action variables. This mechanism enhances the identification of these latent variables over time. Under mild assumptions, we prove that these latent action variables are strictly identifiable. Experimental results on several benchmarks show that the HAL model significantly outperforms existing methods for weakly-supervised action segmentation, confirming its practical effectiveness in real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- ICE-BeeM: Identifiable Conditional Energy-Based Deep Models Based on Nonlinear ICAIlyes Khemakhem, Ricardo Pio Monti, Diederik P. Kingma, Aapo HyvärinenNeurIPS 2020 · 被引用 141 次
- CITRIS: Causal Identifiability from Temporal Intervened SequencesPhillip Lippe, Sara Magliacane, Sindy Löwe, Yuki M. Asano 等ICML 2022 · 被引用 136 次
- Weakly Supervised Energy-Based Learning for Action SegmentationJun Li, Peng Lei, Sinisa TodorovicICCV 2019 · 被引用 109 次
相关 Paper
- Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningZijia Lu, Ehsan ElhamifarICCV 2021 · 被引用 35 次
- Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary AlignmentAngchi Xu, Wei-Shi ZhengCVPR 2024 · 被引用 8 次
- Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention ModelingGuiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang 等ICCV 2023 · 被引用 7 次
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer 等ACL 2020 · 被引用 24 次
- VDSM: Unsupervised Video Disentanglement With State-Space Modeling and Deep Mixtures of ExpertsMatthew J. Vowels, Necati Cihan Camgöz, Richard BowdenCVPR 2021
