Differentiable Grammars for Videos
A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo
摘要
This paper proposes a novel algorithm which learns a formal regular grammar from real-world continuous data, such as videos. Learning latent terminals, non-terminals, and production rules directly from continuous data allows the construction of a generative model capturing sequential structures with multiple possibilities. Our model is fully differentiable, and provides easily interpretable results which are important in order to understand the learned structures. It outperforms the state-of-the-art on several challenging datasets and is more accurate for forecasting future activities in videos. We plan to open-source the code. 1 Learning a formal grammar from continuous, unstructured data is a challenging problem. This is especially challenging when the elements (i.e., terminals) of the grammar to be learned are not symbolic or discrete (Chomsky 1956; 1959) , but are higher dimensional vectors, such as representations from real world data sequences such as videos. Simultaneously, addressing such challenges is necessary for better automated understanding of sequential data. In video understanding, such as activity detection, a convolutional neural network (CNN) (e.g., (Carreira and Zisserman 2017)) generates a representation abstracting local spatiotemporal information at every time step, forming a temporal sequence of representations. Learning a grammar reflecting sequential changes in video representations will enable explicit and high-level modeling of temporal structure and relationships between multiple occurring events in videos. In this paper, we propose a new approach of modeling a formal grammar for videos in terms of learnable and differentiable neural network functions. The objective is to formulate not only the terminals and non-terminals of our grammar as learnable representations but also the production rules, which are generated here as differentiable functions. We provide the loss function to train our differentiable grammar directly from data, and present methodologies to take advantage of it for recognizing and forecasting sequences 2 . Rather
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak 等NeurIPS 2023 · 被引用 17 次
- How Do You Do It? Fine-Grained Action Understanding with Pseudo-AdverbsHazel Doughty, Cees G. M. SnoekCVPR 2022 · 被引用 16 次
- Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsJungin Park, Jiyoung Lee, Kwanghoon SohnCVPR 2025
- Modeling Multi-Label Action Dependencies for Temporal Action LocalizationPraveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahCVPR 2021
相关 Paper
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 被引用 4 次
- Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video DataKyoung-Woon On, Eun-Sol Kim, Yu-Jung Heo, Byoung-Tak ZhangAAAI 2020 · 被引用 7 次
- STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionXi Ye, Guillaume-Alexandre BilodeauAAAI 2024 · 被引用 20 次
- VideoWorld: Exploring Knowledge Learning from Unlabeled VideosZhongwei Ren, Yunchao Wei, Xun Guo, Yao Zhao 等CVPR 2025
- Convolutional Tensor-Train LSTM for Spatio-Temporal LearningJiahao Su, Wonmin Byeon, Jean Kossaifi, Furong Huang 等NeurIPS 2020 · 被引用 146 次
