The Pervasive Blind Spot: Benchmarking VLM Inference Risks on Everyday Personal Videos
Shuning Zhang, Zhaoxin Li, Changxi Wen, Ying Ma, Simin Li, Gengrui Zhang, Ziyi Zhang, Yibo Meng, Hantao Zhao, Xin Yi, Hewu Li
摘要
Applying Vision-Language Models (VLMs) to pervasive personal videos introduces profound privacy risks. This paper addresses the critical yet unexplored inferential privacy threat, specifically the risk of inferring sensitive personal attributes from seemingly benign data. To address this gap, we crowdsourced a dataset of 508 everyday personal videos from 58 individuals, and benchmarked VLM inference capabilities against human performance. Our findings reveal three key insights: (1) VLMs surpass recruited human evaluators in inferential accuracy, analyzing temporal behavioral patterns rather than relying solely on object recognition. (2) Inferential risk is strongly correlated with specific video characteristics and prompting strategies. (3) VLM-driven explanation towards the inference is unreliable, as we observe a disconnect between the model's reasoning and evidential impact, where ubiquitous objects often serve as misleading confounders.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel 等UbiComp 2024 · 被引用 281 次
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 被引用 211 次
相关 Paper
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 被引用 54 次
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang 等ICML 2026 · 被引用 3 次
- Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays 等ICLR 2026 · 被引用 1 次
- The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic FrameworkFeiran Liu, Yuzhe Zhang, Xinyi Huang, Yinan Peng 等ACM MM 2025 · 被引用 4 次
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu 等ICLR 2026 · 被引用 8 次
