Compositional Video Understanding with Spatiotemporal Structure-based Transformers
Hoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol Kim
摘要
In this paper, we suggest a new novel method to understand complex semantic structures through long video inputs. Conventional methods for understanding videos have been focused on short-term clips, and trained to get visual representations for the short clips using convolutional neural networks or transformer architectures. However, most real-world videos are composed of long videos ranging from minutes to hours, therefore, it essentially brings limitations to understanding the overall semantic structures of the long videos by dividing them into small clips and learning the representations of them. We suggest a new algorithm to learn the multi-granular semantic structures of videos, by defining spatiotemporal high-order relationships among object-based representations as semantic units. The proposed method includes a new transformer architecture capable of learning spatiotemporal graphs, and a compositional learning method to learn disentangled features for each semantic unit. Using the suggested method, we resolve the challenging video task, which is compositional generalization understanding of unseen videos. In experiments, we demonstrate new state-of-the-art performances for two challenging video datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Zero-Shot Compositional Video Learning with Coding Rate ReductionHeeseok Jung, Jun-Hyeon Bak, Yujin Jeong, Gyugeun Lee 等ICCV 2025 · 被引用 1 次
- STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-TrainingHaiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan 等CVPR 2025
它引用的顶会 Paper14
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
- Inductive representation learning on temporal graphsDa Xu, Chuanwei Ruan, Evren Körpeoglu, Sushant Kumar 等ICLR 2020 · 被引用 901 次
相关 Paper
- Understanding Video Transformers via Universal Concept DiscoveryMatthew Kowal, Achal Dave, Rares Ambrus, Adrien Gaidon 等CVPR 2024
- Generative Video Transformer: Can Objects be the Words?Yi-Fu Wu, Jaesik Yoon, Sungjin AhnICML 2021 · 被引用 37 次
- Multiview Transformers for Video RecognitionShen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu 等CVPR 2022 · 被引用 279 次
- Mavors: Multi-granularity Video Representation for Multimodal Large Language ModelYang Shi, Jiaheng Liu, Yushuo Guan, Zhenhua Wu 等ACM MM 2025 · 被引用 1 次
- Revisiting Hierarchical Approach for Persistent Long-Term Video PredictionWonkwang Lee, Whie Jung, Han Zhang, Ting Chen 等ICLR 2021 · 被引用 29 次
