USENIX Security2023Top-tier venue
Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
Wenjun Qiu, David Lie, Lisa M. Austin
Abstract
A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric , which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grain labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly 1.71 per labeled text segment. Calpric 's training process also generates a labeled data set of 16K privacy policy text segments across 9 Data categories with balanced positive and negative samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b85a9066-704b-43e2-8d1b-eb382da254f5Cited by top-tier papers2
- Breaking the Illusion: Automated Reasoning of GDPR Consent ViolationsYing Li, Wenjun Qiu, Faysal Hossain Shezan, Kunlin Cai et al.S&P 2026 · 1 citation
- Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated ApproachQinge Xie, Karthik Ramakrishnan, Frank LiUSENIX Security 2025
Builds on6
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub et al.USENIX Security 2018 · 400 citations
- Automated Analysis of Privacy Requirements for Mobile AppsSebastian Zimmeck, Ziqi Wang, Lieyong Zou, Roger Iyengar et al.NDSS 2017 · 255 citations
- PolicyLint: Investigating Internal Privacy Policy Contradictions on Google PlayBenjamin Andow, Samin Yaseer Mahmud, Wenyu Wang, Justin Whitaker et al.USENIX Security 2019 · 185 citations
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch et al.EMNLP 2020 · 144 citations
- Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesShayok ChakrabortyAAAI 2020 · 23 citations
Related papers
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- CrowdRL: An End-to-End Reinforcement Learning Framework for Data LabellingKaiyu Li, Guoliang Li, Yong Wang, Yan Huang et al.ICDE 2021 · 17 citations
- COCA: Cost-Effective Collaborative Annotation System by Combining Experts and AmateursJiayu Lei, Zheng Zhang, Lan Zhang, Xiang-Yang LiICDE 2022 · 6 citations
- PolicyPulse: Precision Semantic Role Extraction for Enhanced Privacy Policy ComprehensionAndrick Adhikari, Sanchari Das, Rinku DewriNDSS 2025
- Intent Classification and Slot Filling for Privacy PoliciesWasi Uddin Ahmad, Jianfeng Chi, Tu Le, Thomas Norton et al.ACL 2021
