SOAR: Scene-debiasing Open-set Action Recognition
Yuanhao Zhai, Ziyi Liu, Zhenyu Wu, Yi Wu, Chunluan Zhou, David S. Doermann, Junsong Yuan, Gang Hua
Abstract
Deep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the open-set action recognition performance when the testing samples have different scene distributions from the training samples. To mitigate this problem, we propose a novel method, called Scene-debiasing Open-set Action Recognition (SOAR), which features an adversarial scene reconstruction module and an adaptive adversarial scene classification module. The former prevents the decoder from reconstructing the video background given video features, and thus helps reduce the background information in feature learning. The latter aims to confuse scene type classification given video features, with a specific emphasis on the action foreground, and helps to learn scene-invariant information. In addition, we design an experiment to quantify the scene bias. The results indicate that the current open-set action recognizers are biased toward the scene, and our proposed SOAR method better mitigates such bias. Furthermore, our extensive experiments demonstrate that our method outperforms state-of-the-art methods, and the ablation studies confirm the effectiveness of our proposed modules.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79aeb249-9f70-4d9d-af1e-820a2fcd819bCited by top-tier papers4
- ContextHOI: Spatial Context Learning for Human-Object Interaction DetectionMingda Jia, Liming Zhao, Ge Li, Yun ZhengAAAI 2025 · 2 citations
- Uncertainty-aware Action Decoupling Transformer for Action AnticipationHongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee et al.CVPR 2024
- Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video BenchmarksNina Shvetsova, Arsha Nagrani, Bernt Schiele, Hilde Kuehne et al.CVPR 2025
- ALBAR: Adversarial Learning approach to mitigate Biases in Action RecognitionJoseph Fioresi, Ishan Rajendrakumar Dave, Mubarak ShahICLR 2025
Builds on21
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Uncertainty Estimation Using a Single Deep Deterministic Neural NetworkJoost van Amersfoort, Lewis Smith, Yee Whye Teh, Yarin GalICML 2020 · 529 citations
- Learning De-biased Representations with Biased RepresentationsHyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo et al.ICML 2020 · 332 citations
- Evidential Deep Learning for Open Set Action RecognitionWentao Bao, Qi Yu, Yu KongICCV 2021 · 204 citations
Related papers
- Learning Discriminative Feature Representation for Open Set Action RecognitionHongjie Zhang, Yi Liu, Yali Wang, Limin Wang et al.ACM MM 2023 · 10 citations
- Mitigating and Evaluating Static Bias of Action Representations in the Background and the ForegroundHaoxin Li, Yuan Liu, Hanwang Zhang, Boyang LiICCV 2023 · 30 citations
- Contrastive Open Set RecognitionBaile Xu, Furao Shen, Jian ZhaoAAAI 2023 · 37 citations
- Generative-Discriminative Feature Representations for Open-Set RecognitionPramuditha Perera, Vlad I. Morariu, Rajiv Jain, Varun Manjunatha et al.CVPR 2020
- Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the MotionJinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu et al.AAAI 2021 · 70 citations
