Gradient Inversion of Multimodal Models
Omri Ben Hemo, Alon Zolfi, Oryan Yehezkel, Omer Hofman, Roman Vainshtein, Hisashi Kojima, Yuval Elovici, Asaf Shabtai
Abstract
Federated learning (FL) enables privacypreserving distributed machine learning by sharing gradients instead of raw data. However, FL remains vulnerable to gradient inversion attacks, in which shared gradients can reveal sensitive training data. Prior research has mainly concentrated on unimodal tasks, particularly image classification, examining the reconstruction of single-modality data, and analyzing privacy vulnerabilities in these relatively simple scenarios. As multimodal models are increasingly used to address complex vision-language tasks, it becomes essential to assess the privacy risks inherent in these architectures. In this paper, we explore gradient inversion attacks targeting multimodal vision-language Document Visual Question Answering (DQA) models and propose GI-DQA, a novel method that reconstructs private document content from gradients. Through extensive evaluation on state-of-the-art DQA models, our approach exposes critical privacy vulnerabilities and highlights the urgent need for robust defenses to secure multimodal FL systems. Project page at: https://AlonZolfi.github.io/GI-DQA/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddc9892d-b78f-4e92-bacf-261a8a6602daBuilds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Robbing the Fed: Directly Obtaining Private Data in Federated Learning with Modified ModelsLiam H. Fowl, Jonas Geiping, Wojciech Czaja, Micah Goldblum et al.ICLR 2022 · 181 citations
- Recovering Private Text in Federated Learning of Language ModelsSamyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao et al.NeurIPS 2022 · 120 citations
Related papers
- Geminio: Language-Guided Gradient Inversion Attacks in Federated LearningJunjie Shan, Ziqi Zhao, Jialin Lu, Rui Zhang et al.ICCV 2025 · 2 citations
- Soteria: Provable Defense Against Privacy Leakage in Federated Learning From Representation PerspectiveJingwei Sun, Ang Li, Binghui Wang, Huanrui Yang et al.CVPR 2021
- Fast Generation-Based Gradient Leakage Attacks against Highly Compressed GradientsDongyun Xue, Haomiao Yang, Mengyu Ge, Jingwei Li et al.INFOCOM 2023 · 4 citations
- SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value DecompositionChenxiang Luo, David K. Y. Yau, Qun SongNDSS 2026 · 3 citations
- SoK: On Gradient Leakage in Federated LearningJiacheng Du, Jiahui Hu, Zhibo Wang, Peng Sun et al.USENIX Security 2025
