Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
Joseph Fioresi, Ishan Rajendrakumar Dave, Mubarak Shah
Abstract
We introduce a novel formulation of visual privacy preservation for video foundation models that operates entirely in the latent space. While spatio-temporal features learned by foundation models have deepened general understanding of video content, sharing or storing these extracted visual features for downstream tasks inadvertently reveals sensitive personal information like skin color, gender, or clothing. Current privacy preservation methods focus on input-pixel-level anonymization, which requires retraining the entire utility video model and results in task-specific anonymization, making them unsuitable for recent video foundational models. To address these challenges, we introduce a lightweight Anonymizing Adapter Module (AAM) that removes private information from video features while retaining general task utility. AAM can be applied in a plug-and-play fashion to frozen video encoders, minimizing the computational burden of finetuning and re-extracting features. Our framework employs three newly designed training objectives: (1) a clip-level self-supervised privacy objective to reduce mutual information between static clips, (2) a co-training objective to retain utility across seen tasks, and (3) a latent consistency loss for generalization on unseen tasks. Our extensive evaluations demonstrate a significant 35% reduction in privacy leakage while maintaining near-baseline utility performance across various downstream tasks: Action Recognition (Kinetics400, UCF101, HMDB51), Temporal Action Detection (THUMOS14), and Anomaly Detection (UCF-Crime). We also provide an analysis on anonymization for sensitive temporal attribute recognition. Additionally, we propose new protocols for assessing gender bias in action recognition models, showing that our method effectively mitigates such biases and promotes more equitable video understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Differentially Private Fine-tuning of Language ModelsDa Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi et al.ICLR 2022 · 494 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Toyota Smarthome: Real-World Activities of Daily LivingSrijan Das, Rui Dai, Michal Koperski, Luca Minciullo et al.ICCV 2019 · 182 citations
- FACET: Fairness in Computer Vision Evaluation BenchmarkLaura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval et al.ICCV 2023 · 74 citations
Related papers
- TeD-SPAD: Temporal Distinctiveness for Self-supervised Privacy-preservation for video Anomaly DetectionJoseph Fioresi, Ishan Rajendrakumar Dave, Mubarak ShahICCV 2023 · 35 citations
- STPrivacy: Spatio-Temporal Privacy-Preserving Action RecognitionMing Li, Xiangyu Xu, Hehe Fan, Pan Zhou et al.ICCV 2023 · 40 citations
- SPAct: Self-supervised Privacy Preservation for Action RecognitionIshan Rajendrakumar Dave, Chen Chen, Mubarak ShahCVPR 2022 · 62 citations
- Privacy-Aware Video Anomaly Detection: Guided Orthogonal Projection and a Comprehensive Evaluation FrameworkWenxiang Diao, Lei Wang, Andrew Busch, Jun Zhou et al.ICML 2026
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang et al.ICLR 2023 · 62 citations
