JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity Detection
Mahsa Ehsanpour, Fatemeh Sadat Saleh, Silvio Savarese, Ian D. Reid, Hamid Rezatofighi
摘要
The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, with potentially highly unbalanced and longtailed distributed action labels from a stream of sensory data captured from a mobile robot platform remains a significant challenge, not least owing to the lack of a reflective large-scale dataset. In this paper, we introduce JRDB-Act, as an extension of the existing JRDB, which is captured by a social mobile manipulator and reflects a real distribution of human daily-life actions in a university campus environment. JRDB-Act has been densely annotated with atomic actions, comprises over 2.8M action labels, constituting a large-scale spatio-temporal action detection dataset. Each human bounding box is labeled with one pose-based action label and multiple (optional) interaction-based action labels. Moreover JRDB-Act provides social group annotation, conducive to the task of grouping individuals based on their interactions in the scene to infer their social activities (common activities in each social group). Each annotated label in JRDB-Act is tagged with the annotators' confidence level which contributes to the development of reliable evaluation strategies. In order to demonstrate how one can effectively utilise such annotations, we develop an end-to-end trainable pipeline to learn and infer these tasks, i.e. individual action and social group detection. The data and the evaluation code will be publicly available at https://jrdb.erc.monash.edu/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Chaotic World: A Large and Challenging Benchmark for Human Behavior Understanding in Chaotic EventsKian Eng Ong, Xun Long Ng, Yanchao Li, Wenjie Ai 等ICCV 2023 · 被引用 6 次
- AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity RecognitionMeiqi Cao, Rui Yan, Xiangbo Shu, Guangzhao Dai 等ACM MM 2024 · 被引用 3 次
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in RoboticsSimindokht Jahangard, Mehrzad Mohammadi, Yi Shen, Zhixi Cai 等AAAI 2026 · 被引用 2 次
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction DetectionDongkeun Kim, Minsu Cho, Suha KwakNeurIPS 2025 · 被引用 1 次
- Dynamic Group Detection using VLM-augmented Temporal Groupness GraphKaname Yokoyama, Chihiro Nakatani, Norimichi UkitaICCV 2025 · 被引用 1 次
它引用的顶会 Paper4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerShuaicheng Li, Qianggang Cao, Lingbo Liu, Kunlin Yang 等ICCV 2021 · 被引用 149 次
- Overcoming Classifier Imbalance for Long-Tail Object Detection With Balanced Group SoftmaxYu Li, Tao Wang, Bingyi Kang, Sheng Tang 等CVPR 2020
相关 Paper
- JRDB-Pose: A Large-Scale Dataset for Multi-Person Pose Estimation and TrackingEdward Vendrow, Duy-Tho Le, Jianfei Cai, Hamid RezatofighiCVPR 2023
- Recognizing Actions From Robotic View for Natural Human-Robot InteractionZiyi Wang, Peiming Li, Hong Liu, Zhichao Deng 等ICCV 2025 · 被引用 1 次
- BABEL: Bodies, Action and Behavior With English LabelsAbhinanda R. Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez 等CVPR 2021
- MMAD: Multi-Label Micro-Action Detection in VideosKun Li, Pengyu Liu, Dan Guo, Fei Wang 等ICCV 2025 · 被引用 21 次
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 等CVPR 2020
