Mitigating and Evaluating Static Bias of Action Representations in the Background and the Foreground
Haoxin Li, Yuan Liu, Hanwang Zhang, Boyang Li
摘要
In video action recognition, shortcut static features can interfere with the learning of motion features, resulting in poor out-of-distribution (OOD) generalization. The video background is clearly a source of static bias, but the video foreground, such as the clothing of the actor, can also provide static bias. In this paper, we empirically verify the existence of foreground static bias by creating test videos with conflicting signals from the static and moving portions of the video. To tackle this issue, we propose a simple yet effective technique, StillMix, to learn robust action representations. Specifically, StillMix identifies bias-inducing video frames using a 2D reference network and mixes them with videos for training, serving as effective bias suppression even when we cannot explicitly extract the source of bias within each video frame or enumerate types of bias. Finally, to precisely evaluate static bias, we synthesize two new benchmarks, SCUBA for static cues in the background, and SCUFO for static cues in the foreground. With extensive experiments, we demonstrate that StillMix mitigates both types of static bias and improves video representations for downstream applications. Code is available at https://github.com/lihaoxin05/StillMix .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Diversifying Spatial-Temporal Perception for Video Domain GeneralizationKun-Yu Lin, Jia-Run Du, Yipeng Gao, Jiaming Zhou 等NeurIPS 2023 · 被引用 27 次
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action RecognitionJongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim 等NeurIPS 2025 · 被引用 4 次
- Don't Judge by the Look: Towards Motion Coherent Video RepresentationYitian Zhang, Yue Bai, Huan Wang, Yizhou Wang 等ICLR 2024 · 被引用 3 次
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language ModelsMaya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari 等NeurIPS 2025 · 被引用 3 次
- Debiased Active Learning with Variational Gradient RectifierWeiguo Chen, Changjian Wang, Shijun Li, Kele Xu 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- Motion-aware Contrastive Video Representation Learning via Foreground-background MergingShuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian 等CVPR 2022 · 被引用 54 次
- Self-Supervised Motion Learning From Static ImagesZiyuan Huang, Shiwei Zhang, Jianwen Jiang, Mingqian Tang 等CVPR 2021
- SOAR: Scene-debiasing Open-set Action RecognitionYuanhao Zhai, Ziyi Liu, Zhenyu Wu, Yi Wu 等ICCV 2023 · 被引用 15 次
- Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation LearningManlin Zhang, Jinpeng Wang, Andy J. MaAAAI 2022 · 被引用 9 次
- Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation LearningJinpeng Wang, Yuting Gao, Ke Li, Yiqi Lin 等CVPR 2021
