Ordered Atomic Activity for Fine-grained Interactive Traffic Scenario Understanding
Nakul Agarwal, Yi-Ting Chen
Abstract
We introduce a novel representation called Ordered Atomic Activity for interactive scenario understanding. The representation decomposes each scenario into a set of ordered atomic activities, where each activity consists of an action and the corresponding actors involved and the order denotes the temporal development of the scenario. This design also helps in identifying important interactive relationships, such as yielding. The action is a high-level semantic motion pattern that is grounded in the surrounding road topology, which we decompose into zones and corners with unique IDs. For example, a group of pedestrians crossing in front is denoted as C1 → C4: P+, as depicted in Figure 1 . We collect a new large-scale dataset called OATS 1 (Ordered Atomic Activities in interactive Traffic Scenarios), comprising 1026 video clips (∼ 20s) captured at intersections in San Francisco Bay Area. Each clip is labeled with the proposed language, resulting in 59 activity categories and 6512 annotated activity instances. We propose three fine-grained scenario understanding tasks, i.e., multilabel atomic activity recognition, activity order prediction, and interactive scenario retrieval. We also propose a Graph Convolutional Network based framework that models both appearance and motion of traffic participants to tackle the above tasks, that performs favorably against state-of-theart methods. However, we find that the methods cannot achieve satisfactory performance, indicating rising opportunities for the community to develop new algorithms for these tasks towards better interactive scenario understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af4fb99e-dea7-4282-ae57-58554751e70cCited by top-tier papers2
- TUMTraf VideoQA: Dataset and Benchmark for Unified Spatio-Temporal Video Understanding in Traffic ScenesXingcheng Zhou, Konstantinos Larintzakis, Hao Guo, Walter Zimmer et al.ICML 2025
- Action-Slot: Visual Action-Centric Representations for Multi-Label Atomic Activity Recognition in Traffic ScenesChi-Hsi Kung, Shu-Wei Lu, Yi-Hsuan Tsai, Yi-Ting ChenCVPR 2024
Builds on16
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory PredictionAmir Rasouli, Iuliia Kotseruba, Toni Kunic, John K. TsotsosICCV 2019 · 411 citations
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular CamerasAnthony Hu, Zak Murez, Nikhil Mohan, Sofía Dudas et al.ICCV 2021 · 329 citations
- WoodScape: A Multi-Task, Multi-Camera Fisheye Dataset for Autonomous DrivingSenthil Kumar Yogamani, Christian Witt, Hazem Rashed, Sanjaya Nayak et al.ICCV 2019 · 325 citations
Related papers
- HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video UnderstandingTrong-Thuan Nguyen, Pha A. Nguyen, Khoa LuuCVPR 2024 · 5 citations
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
- MOMA: Multi-Object Multi-Actor Activity ParsingZelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang et al.NeurIPS 2021 · 34 citations
- Understanding Human Gaze Communication by Spatio-Temporal Graph ReasoningLifeng Fan, Wenguan Wang, Song-Chun Zhu, Xinyu Tang et al.ICCV 2019 · 124 citations
- Action Scene Graphs for Long-Form Understanding of Egocentric VideosIvan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi et al.CVPR 2024 · 14 citations
