VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
Xiaobin Hu, Enpu zuo, Lanping Hu, Kaiwen Yang, Dianshu Liao, Tianyi Zhang, Bo Yin, Yinsi Zhou, Shidong Pan, xiaoyu sun
Abstract
Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust privacy detection algorithms. However, current robust detection models are severely hindered by the lack of comprehensive datasets. Existing privacy-oriented datasets often suffer from limited scale, coarse-grained annotations, and narrow domain coverage, failing to capture the intricate details of sensitive information in real-world environments. To bridge this gap, we present a large-scale, fine-grained Visual Privacy Dataset (VPD-100K), designed to facilitate generalized privacy detection. We establish a holistic taxonomy comprising four primary domains: Human Presence, On-Screen Personally Identifiable Information (PII), Physical Identifiers, and Location Indicators, containing 100,000 images annotated with 33 fine-grained classes and over 190,000 object instances. Statistical analysis reveals that our dataset features long-tailed distributions, small object scales, and high visual complexity. These characteristics make the dataset particularly valuable for demanding, unconstrained applications such as live streaming, where actors frequently face unintentional, real-time information leakage. Furthermore, we design an effective frequency-enhance lightweight module consisting of frequency-domain attention fusion and adaptive spectral gating mechanism that breaks the limitations of spatial pixel intensity to better capture the subtle details of sensitive information. Extensive experiments conducted on both diverse image and streaming videos benchmarks consistently demonstrate the effectiveness of our VPD-100K dataset and the well-curated frequency mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd010671-2671-409b-ae37-8d3bc55655c0Builds on8
- Gold-YOLO: Efficient Object Detector via Gather-and-Distribute MechanismChengcheng Wang, Wei He, Ying Nie, Jianyuan Guo et al.NeurIPS 2023 · 732 citations
- FBRT-YOLO: Faster and Better for Real-Time Aerial Image DetectionYao Xiao, Tingfa Xu, Yu Xin, Jianan LiAAAI 2025 · 115 citations
- Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected With People Who Are BlindTanusree Sharma, Abigale Stangl, Lotus Zhang, Yu-Yun Tseng et al.CHI 2023 · 23 citations
- Do Streamers Care about Bystanders' Privacy? An Examination of Live Streamers' Considerations and Strategies for Bystanders' Privacy ManagementYanlai Wu, Xinning Gui, Pamela J. Wisniewski, Yao LiCSCW 2023 · 15 citations
- DIPA2: An Image Dataset with Cross-cultural Privacy Perception AnnotationsAnran Xu, Zhongyi Zhou, Kakeru Miyazaki, Ryo Yoshikawa et al.UbiComp 2024 · 15 citations
Related papers
- Large-scale Video Panoptic Segmentation in the Wild: A BenchmarkJiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li et al.CVPR 2022 · 58 citations
- Towards Real-World Prohibited Item Detection: A Large-Scale X-ray BenchmarkBoying Wang, Libo Zhang, Longyin Wen, Xianglong Liu et al.ICCV 2021 · 110 citations
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo et al.CVPR 2020
- Triple-Cooperative Video Shadow DetectionZhihao Chen, Liang Wan, Lei Zhu, Jia Shen et al.CVPR 2021
- Characterizing and Detecting Non-Consensual Photo Sharing on Social NetworksTengfei Zheng, Tongqing Zhou, Qiang Liu, Kui Wu et al.CCS 2022 · 6 citations
