Action-Slot: Visual Action-Centric Representations for Multi-Label Atomic Activity Recognition in Traffic Scenes
Chi-Hsi Kung, Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting Chen
Abstract
Figure 1 . Illustration of the concept of multi-label atomic activity recognition and our proposed Action-slot. In the scene, three atomic activities are presented and depicted by colored arrows. For example, the red arrow represents the Z1-Z4: C+ atomic activity, indicating a group of vehicles turning left. Atomic activities are defined based on road user's type and their motion patterns grounded in the underlying road structure. We introduce Action-slot to learn visual action-centric representations that enable decomposing multiple atomic activities in videos. We demonstrate that our framework can effectively recognize multiple atomic activities via learned representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d348b427-2eed-4865-b942-a862877bd87dCited by top-tier papers4
- What Changed and What Could Have Changed? State-Change Counterfactuals for Procedure-Aware Video Representation LearningChi-Hsi Kung, Frangil Ramirez, Juhyung Ha, Yi-Ting Chen et al.ICCV 2025 · 3 citations
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric PerspectiveNhat Chung, Taisei Hanyu, Toan Nguyen, Huy Le et al.AAAI 2026 · 1 citation
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer et al.ICML 2025
- PatchDEMUX: A Certifiably Robust Framework for Multi-label Classifiers Against Adversarial PatchesDennis Jacob, Chong Xiang, Prateek MittalCVPR 2025
Builds on35
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Ordered Atomic Activity for Fine-grained Interactive Traffic Scenario UnderstandingNakul Agarwal, Yi-Ting ChenICCV 2023 · 7 citations
- MOMA: Multi-Object Multi-Actor Activity ParsingZelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang et al.NeurIPS 2021 · 34 citations
- Home Action Genome: Cooperative Compositional Action UnderstandingNishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai et al.CVPR 2021
- Multi-Label Activity Recognition Using Activity-Specific Features and Activity CorrelationsYanyi Zhang, Xinyu Li, Ivan MarsicCVPR 2021
- Shepherding Slots to Objects: Towards Stable and Robust Object-Centric LearningJinwoo Kim, Janghyuk Choi, Ho-Jin Choi, Seon Joo KimCVPR 2023
