Data Guard: A Fine-Grained Purpose-Based Access Control System for Large Data Warehouses
Khai Tran, Sudarshan Vasudevan, Pratham Desai, Alex Gorelik, Mayank Ahuja, Athrey Yadatore Venkateshababu, Mohit Verma, Dichao Hu, Walaa Eldin Moustafa, Vasanth Rajamani, Ankit Gupta, Issac Buenrostro
摘要
The last few years have witnessed a spate of data protection regulations in conjunction with an ever-growing appetite for data usage in large businesses, which presents significant challenges for businesses to maintain compliance. To address this conflict, we present Data Guard - a fine-grained, purpose-based access control system for large data warehouses. Data Guard enables authoring policies based on semantic descriptions of data and purpose of data access. Data Guard then translates these policies into SQL views that mask data from the underlying warehouse tables. At access time, Data Guard ensures compliance by transparently routing each table access to the appropriate data-masking view based on the purpose of the access, thus minimizing the effort of adopting Data Guard in existing applications. Our enforcement solution allows masking data at much finer granularities than what traditional solutions allow. In addition to row and column level data masking, Data Guard can mask data at the sub-cell level for columns with non-atomic data types such as structs, arrays, and maps. This fine-grained masking allows Data Guard to preserve data utility for consumers while ensuring compliance. We implemented a number of performance optimizations to minimize the overhead of data masking operations. We perform numerous experiments to identify the key factors that influence the data masking overhead and demonstrate the efficiency of our implementation. Data Guard is deployed inside LinkedIn's production data warehouses and ensures compliance of more than 20,000 table accesses each day across different data processing engines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Thoth: Comprehensive Policy Compliance in Data Retrieval SystemsEslam Elnikety, Aastha Mehta, Anjo Vahldiek-Oberwagner, Deepak Garg 等USENIX Security 2016 · 被引用 30 次
- Disclosure-Compliant Query AnsweringRudi Poepsel Lemaitre, Kaustubh Beedkar, Volker MarklSIGMOD 2025 · 被引用 1 次
- TaintStream: fine-grained taint tracking for big data platforms through dynamic code translationChengxu Yang, Yuanchun Li, Mengwei Xu, Zhenpeng Chen 等FSE 2021 · 被引用 11 次
- Software-Defined Data Protection: Low Overhead Policy Compliance at the Storage Layer is Within Reach!Zsolt István, Soujanya Ponnapalli, Vijay ChidambaramVLDB 2021 · 被引用 19 次
- The Taming of the Stack: Isolating Stack Data from Memory ErrorsKaiming Huang, Yongzhe Huang, Mathias Payer, Zhiyun Qian 等NDSS 2022
