OPEL: Optimal Transport Guided ProcedurE Learning
Sayeed Shafayet Chowdhury, Soumyadeep Chandra, Kaushik Roy
摘要
Procedure learning refers to the task of identifying the key-steps and determining their logical order, given several videos of the same task. For both third-person and first-person (egocentric) videos, state-of-the-art (SOTA) methods aim at finding correspondences across videos in time to accomplish procedure learning. However, to establish temporal relationships within the sequences, these methods often rely on frame-to-frame mapping, or assume monotonic alignment of video pairs, leading to sub-optimal results. To this end, we propose to treat the video frames as samples from an unknown distribution, enabling us to frame their distance calculation as an optimal transport (OT) problem. Notably, the OT-based formulation allows us to relax the previously mentioned assumptions. To further improve performance, we enhance the OT formulation by introducing two regularization terms. The first, inverse difference moment regularization, promotes transportation between instances that are homogeneous in the embedding space as well as being temporally closer. The second, regularization based on the KL-divergence with an exponentially decaying prior smooths the alignment while enforcing conformity to the optimality (alignment obtained from vanilla OT optimization) and temporal priors. The resultant optimal transport guided procedure learning framework (‘OPEL’) significantly outperforms the SOTA on benchmark datasets. Specifically, we achieve 22.4% (IoU) and 26.9% (F1) average improvement compared to the current SOTA on large scale egocentric benchmark, EgoProceL. Furthermore, for the third person benchmarks (ProCeL and CrossTask), the proposed approach obtains 46.2% (F1) average enhancement over SOTA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Joint Self-Supervised Video Alignment and Action SegmentationAli Shah Ali, Syed Ahmed Mahmood, Mubin Saeed, Andrey Konin 等ICCV 2025 · 被引用 13 次
- HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric VideosSimone Alberto Peirone, Francesca Pistilli, Giuseppe AvertaICCV 2025
- Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal TransportYang Xiao, Weiming Liu, Jun Dan, Tengyue Xu 等CVPR 2026
它引用的顶会 Paper16
- DynamoNet: Dynamic Action and Motion NetworkAli Diba, Vivek Sharma, Luc Van Gool, Rainer StiefelhagenICCV 2019 · 被引用 123 次
- Unsupervised Procedure Learning via Joint Dynamic SummarizationEhsan Elhamifar, Zwe NaingICCV 2019 · 被引用 61 次
- Unsupervised Action Segmentation by Joint Representation Learning and Online ClusteringSateesh Kumar, Sanjay Haresh, Awais Ahmed, Andrey Konin 等CVPR 2022 · 被引用 52 次
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer 等ACL 2020 · 被引用 24 次
- P3IV: Probabilistic Procedure Planning from Instructional Videos with Weak SupervisionHe Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G. Derpanis 等CVPR 2022 · 被引用 23 次
相关 Paper
- Learning to Align Sequential Actions in the WildWeizhe Liu, Bugra Tekin, Huseyin Coskun, Vibhav Vineet 等CVPR 2022
- De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningChong Liu, Xun Jiang, Fumin Shen, Lei Zhu 等AAAI 2026
- EgoTV: Egocentric Task Verification from Natural Language Task DescriptionsRishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra 等ICCV 2023
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric VideoYuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li 等CVPR 2026 · 被引用 1 次
- Learning by Aligning Videos in TimeSanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram Najam Syed 等CVPR 2021
