Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
Wenjun Qiu, David Lie, Lisa M. Austin
摘要
A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric , which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grain labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly 1.71 per labeled text segment. Calpric 's training process also generates a labeled data set of 16K privacy policy text segments across 9 Data categories with balanced positive and negative samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Breaking the Illusion: Automated Reasoning of GDPR Consent ViolationsYing Li, Wenjun Qiu, Faysal Hossain Shezan, Kunlin Cai 等S&P 2026 · 被引用 1 次
- Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated ApproachQinge Xie, Karthik Ramakrishnan, Frank LiUSENIX Security 2025
它引用的顶会 Paper6
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub 等USENIX Security 2018 · 被引用 400 次
- Automated Analysis of Privacy Requirements for Mobile AppsSebastian Zimmeck, Ziqi Wang, Lieyong Zou, Roger Iyengar 等NDSS 2017 · 被引用 255 次
- PolicyLint: Investigating Internal Privacy Policy Contradictions on Google PlayBenjamin Andow, Samin Yaseer Mahmud, Wenyu Wang, Justin Whitaker 等USENIX Security 2019 · 被引用 185 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesShayok ChakrabortyAAAI 2020 · 被引用 23 次
相关 Paper
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- CrowdRL: An End-to-End Reinforcement Learning Framework for Data LabellingKaiyu Li, Guoliang Li, Yong Wang, Yan Huang 等ICDE 2021 · 被引用 17 次
- COCA: Cost-Effective Collaborative Annotation System by Combining Experts and AmateursJiayu Lei, Zheng Zhang, Lan Zhang, Xiang-Yang LiICDE 2022 · 被引用 6 次
- PolicyPulse: Precision Semantic Role Extraction for Enhanced Privacy Policy ComprehensionAndrick Adhikari, Sanchari Das, Rinku DewriNDSS 2025
- Intent Classification and Slot Filling for Privacy PoliciesWasi Uddin Ahmad, Jianfeng Chi, Tu Le, Thomas Norton 等ACL 2021
