Bridge-Prompt: Towards Ordinal Action Understanding in Instructional Videos
Muheng Li, Lei Chen, Yueqi Duan, Zhilan Hu, Jianjiang Feng, Jie Zhou, Jiwen Lu
摘要
Action recognition models have shown a promising capability to classify human actions in short video clips. In a real scenario, multiple correlated human actions commonly occur in particular orders, forming semantically meaningful human activities. Conventional action recognition approaches focus on analyzing single actions. However, they fail to fully reason about the contextual relations between adjacent actions, which provide potential temporal logic for understanding long videos. In this paper, we propose a prompt-based framework, Bridge-Prompt (Br-Prompt), to model the semantics across adjacent actions, so that it simultaneously exploits both out-of-context and contextual information from a series of ordinal actions in instructional videos. More specifically, we reformulate the individual action labels as integrated text prompts for supervision, which bridge the gap between individual action semantics. The generated text prompts are paired with corresponding video clips, and together co-train the text encoder and the video encoder via a contrastive approach. The learned vision encoder has a stronger capability for ordinal-action-related downstream tasks, e.g. action segmentation and human activity recognition. We evaluate the performances of our approach on several video datasets: Georgia Tech Egocentric Activities (GTEA), 50Salads, and the Breakfast dataset. Br-Prompt achieves state-of-the-art on multiple benchmarks. Code is available at: https: //github.com/ttlmh/Bridge-Prompt .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Diffusion Action SegmentationDaochang Liu, Qiyue Li, Anh-Dung Dinh, Tingting Jiang 等ICCV 2023 · 被引用 113 次
- Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationYuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun 等ICCV 2023 · 被引用 111 次
- PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionCong Wang, Jinshan Pan, Wei Wang, Jiangxin Dong 等NeurIPS 2023 · 被引用 109 次
- SelfPromer: Self-Prompt Dehazing Transformers with Depth-ConsistencyCong Wang, Jinshan Pan, Wanyu Lin, Jiangxin Dong 等AAAI 2024 · 被引用 61 次
- How Much Temporal Long-Term Context is Needed for Action Segmentation?Emad Bahrami Rad, Gianpiero Francesca, Juergen GallICCV 2023 · 被引用 54 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- HierVL: Learning Hierarchical Video-Language EmbeddingsKumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen GraumanCVPR 2023
- Event-Guided Procedure Planning from Instructional Videos with Text SupervisionAn-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng 等ICCV 2023 · 被引用 21 次
- Improving Action Segmentation via Graph-Based Temporal ReasoningYifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2020
- EgoPrompt: Prompt Learning for Egocentric Action RecognitionHuaihai Lyu, Chaofan Chen, Yuheng Ji, Changsheng XuACM MM 2025 · 被引用 3 次
- Open Set Video HOI detection from Action-centric Chain-of-Look PromptingNan Xi, Jingjing Meng, Junsong YuanICCV 2023 · 被引用 10 次
