Toyota Smarthome: Real-World Activities of Daily Living
Srijan Das, Rui Dai, Michal Koperski, Luca Minciullo, Lorenzo Garattoni, François Brémond, Gianpiero Francesca
摘要
The performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce a large real-world video dataset for activities of daily living: Toyota Smarthome. The dataset consists of 16K RGB+D clips of 31 activity classes, performed by seniors in a smarthome. Unlike previous datasets, videos were fully unscripted. As a result, the dataset poses several challenges: high intra-class variation, high class imbalance, simple and composite activities, and activities with similar motion and variable duration. Activities were annotated with both coarse and fine-grained labels. These characteristics differentiate Toyota Smarthome from other datasets for activity recognition. As recent activity recognition approaches fail to address the challenges posed by Toyota Smarthome, we present a novel activity recognition method with attention mechanism. We propose a pose driven spatiotemporal attention mechanism through 3D ConvNets. We show that our novel method outperforms state-of-the-art methods on benchmark datasets, as well as on the Toyota Smarthome dataset. We release the dataset for research use 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 被引用 37 次
- Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art EvaluationMichal Balazia, Philipp Müller, Ákos Levente Tánczos, August von Liechtenstein 等ACM MM 2022 · 被引用 28 次
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 18 次
- Self-Supervised Video Representation Learning via Latent Time NavigationDi Yang, Yaohui Wang, Quan Kong, Antitza Dantcheva 等AAAI 2023 · 被引用 18 次
- A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action RecognitionAndong Deng, Taojiannan Yang, Chen ChenICCV 2023 · 被引用 18 次
它引用的顶会 Paper1
相关 Paper
- Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous VehiclesManuel Martin, Alina Roitberg, Monica Haurilet, Matthias Horne 等ICCV 2019 · 被引用 235 次
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
- Towards Hazardous Activity Recognition for A Novel Real-World DatasetShehzad Ali, Md Tanvir Islam, Ik Hyun Lee, Mingfu Xiong 等ACM MM 2025 · 被引用 1 次
- Recognizing Actions in Videos From Unseen ViewpointsA. J. Piergiovanni, Michael S. RyooCVPR 2021
- JRDB-Act: A Large-scale Dataset for Spatio-temporal Action, Social Group and Activity DetectionMahsa Ehsanpour, Fatemeh Sadat Saleh, Silvio Savarese, Ian D. Reid 等CVPR 2022 · 被引用 49 次
