JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection
Mahsa Ehsanpour, Fatemeh Sadat Saleh, Silvio Savarese, Ian D. Reid, Hamid Rezatofighi
Abstract
The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, with potentially highly unbalanced and longtailed distributed action labels from a stream of sensory data captured from a mobile robot platform remains a significant challenge, not least owing to the lack of a reflective large-scale dataset. In this paper, we introduce JRDB-Act, as an extension of the existing JRDB, which is captured by a social mobile manipulator and reflects a real distribution of human daily-life actions in a university campus environment. JRDB-Act has been densely annotated with atomic actions, comprises over 2.8M action labels, constituting a large-scale spatio-temporal action detection dataset. Each human bounding box is labeled with one pose-based action label and multiple (optional) interaction-based action labels. Moreover JRDB-Act provides social group annotation, conducive to the task of grouping individuals based on their interactions in the scene to infer their social activities (common activities in each social group). Each annotated label in JRDB-Act is tagged with the annotators' confidence level which contributes to the development of reliable evaluation strategies. In order to demonstrate how one can effectively utilise such annotations, we develop an end-to-end trainable pipeline to learn and infer these tasks, i.e. individual action and social group detection. The data and the evaluation code will be publicly available at https://jrdb.erc.monash.edu/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4bc6c62-ff56-4fdd-a8b6-d342f1cc6459Cited by top-tier papers12
- Chaotic World: A Large and Challenging Benchmark for Human Behavior Understanding in Chaotic EventsKian Eng Ong, Xun Long Ng, Yanchao Li, Wenjie Ai et al.ICCV 2023 · 6 citations
- AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity RecognitionMeiqi Cao, Rui Yan, Xiangbo Shu, Guangzhao Dai et al.ACM MM 2024 · 3 citations
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in RoboticsSimindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai et al.AAAI 2026 · 2 citations
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction DetectionDongkeun Kim, Minsu Cho, Suha KwakNeurIPS 2025 · 1 citation
- Dynamic Group Detection using VLM-augmented Temporal Groupness GraphKaname Yokoyama, Chihiro Nakatani, Norimichi UkitaICCV 2025 · 1 citation
Builds on4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang et al.ICCV 2021 · 149 citations
- Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxYu Li, Tao Wang, Bingyi Kang, Sheng Tang et al.CVPR 2020
Related papers
- JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and TrackingEdward Vendrow, Duy-Tho Le, Jianfei Cai, Hamid RezatofighiCVPR 2023
- Recognizing Actions From Robotic View for Natural Human-Robot InteractionZiyi Wang, Peiming Li, Hong Liu, Zhichao Deng et al.ICCV 2025 · 1 citation
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez et al.CVPR 2021
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang et al.ICCV 2025 · 21 citations
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo et al.CVPR 2020
