Can Deep Learning Recognize Subtle Human Activities?
Vincent Jacquot, Zhuofan Ying, Gabriel Kreiman
Abstract
Deep Learning has driven recent and exciting progress in computer vision, instilling the belief that these algorithms could solve any visual task. Yet, datasets commonly used to train and test computer vision algorithms have pervasive confounding factors. Such biases make it difficult to truly estimate the performance of those algorithms and how well computer vision models can extrapolate outside the distribution in which they were trained. In this work, we propose a new action classification challenge that is performed well by humans, but poorly by state-of-the-art Deep Learning models. As a proof-of-principle, we consider three exemplary tasks: drinking, reading, and sitting. The best accuracies reached using state-of-the-art computer vision models were 61.7%, 62.8%, and 76.8%, respectively, while human participants scored above 90% accuracy on the three tasks. We propose a rigorous method to reduce confounds when creating datasets, and when comparing human versus computer vision performance. Source code and datasets are publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1d8feaf-4102-4905-be77-dd3de4c86ad6Cited by top-tier papers2
- Flow Snapshot Neurons in Action: Deep Neural Networks Generalize to Biological Motion PerceptionShuangpeng Han, Ziyu Wang, Mengmi ZhangNeurIPS 2024 · 8 citations
- HumorDB: Can AI Understand Graphical Humor?Veedant Jain, Gabriel Kreiman, Felipe dos Santos Alves FeitosaICCV 2025 · 1 citation
Builds on1
Related papers
- Human and AI Perceptual Differences in Image Classification ErrorsMinghao Liu, Jiaheng Wei, Yang Liu, James DavisAAAI 2025 · 11 citations
- The 3D-PC: a benchmark for visual perspective taking in humans and machinesDrew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj et al.ICLR 2025
- A Decade's Battle on Dataset Bias: Are We There Yet?Zhuang Liu, Kaiming HeICLR 2025 · 8 citations
- Testing DNN image classifiers for confusion & bias errorsYuchi Tian, Ziyuan Zhong, Vicente Ordonez, Gail E. Kaiser et al.ICSE 2020 · 33 citations
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo et al.ICCV 2019 · 182 citations
