Vid-SME: Membership Inference Attacks against Large Video Understanding Models
Qi Li, Runpeng Yu, Xinchao Wang
Abstract
Multimodal large language models (MLLMs) demonstrate remarkable capabilities in handling complex multimodal tasks and are increasingly adopted in video understanding applications. However, their rapid advancement raises serious data privacy concerns, particularly given the potential inclusion of sensitive video content, such as personal recordings and surveillance footage, in their training datasets. Determining improperly used videos during training remains a critical and unresolved challenge. Despite considerable progress on membership inference attacks (MIAs) for text and image data in MLLMs, existing methods fail to generalize effectively to the video domain. These methods suffer from poor scalability as more frames are sampled and generally achieve negligible true positive rates at low false positive rates (TPR@Low FPR), mainly due to their failure to capture the inherent temporal variations of video frames and to account for model behavior differences as the number of frames varies. To address these challenges, we introduce Vid-SME (Video Sharma-Mittal Entropy), the first membership inference method tailored for video data used in video understanding LLMs (VULLMs). Vid-SME leverages the confidence of model output and integrates adaptive parameterization to compute Sharma-Mittal entropy (SME) for video inputs. By leveraging the SME difference between natural and temporally-reversed video frames, Vid-SME derives robust membership scores to determine whether a given video is part of the model's training set. Experiments on various self-trained and open-sourced VULLMs demonstrate the strong effectiveness of Vid-SME. Code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef35eaea-57b1-403c-9c4e-0a52ce31adb8Cited by top-tier papers5
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement LearningSicheng Feng, Kaiwen Tuo, Song Wang, Lingdong Kong et al.ICLR 2026 · 28 citations
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and UnderstandingYan Shu, Hangui Lin, Yexin Liu, Yan Zhang et al.NeurIPS 2025 · 17 citations
- VICTOR: Dataset Copyright Auditing in Video Recognition SystemsQuan Yuan, Zhikun Zhang, Linkang Du, Min Chen et al.NDSS 2026 · 2 citations
- CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset TrainingQi Li, Cheng-Long Wang, Yinzhi Cao, Di WangACL 2026 · 2 citations
- Black-Box Membership Inference Attacks for Video Training Data in Multimodal Large Language ModelsJinrui Wang, Zhenfeng Gao, Wendan Wang, Huili Wang et al.ACL 2026
Builds on28
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- Membership Inference Attacks against Large Vision-Language ModelsZhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin et al.NeurIPS 2024 · 43 citations
- Membership Inference Attacks Against Vision-Language ModelsYuke Hu, Zheng Li, Zhihao Liu, Yang Zhang et al.USENIX Security 2025
- VidLeaks: Membership Inference Attacks Against Text-to-Video ModelsLi Wang, Wenyu Chen, Ning Yu, Zheng Li et al.USENIX Security 2026 · 2 citations
- LOMIA: Label-Only Membership Inference Attacks against Pre-trained Large Vision-Language ModelsYihao Liu, Xinqi Lyu, Dong Wang, Yanjie Li et al.NeurIPS 2025 · 3 citations
- Robust Membership Inference for Large Language Models under Adversarial Generative CorruptionYuanhong Huang, Huili Wang, Xueying Bai, Jinrui Wang et al.ACL 2026
