Hark: A Deep Learning System for Navigating Privacy Feedback at Scale
Hamza Harkous, Sai Teja Peddinti, Rishabh Khandelwal, Animesh Srivastava, Nina Taft
摘要
Integrating user feedback is one of the pillars for building successful products. However, this feedback is generally collected in an unstructured free-text form, which is challenging to understand at scale. This is particularly demanding in the privacy domain due to the nuances associated with the concept and the limited existing solutions. In this work, we present Hark1, a system for discovering and summarizing privacy-related feedback at scale. Hark automates the entire process of summarizing privacy feedback, starting from unstructured text and resulting in a hierarchy of high-level privacy themes and fine-grained issues within each theme, along with representative reviews for each issue. At the core of Hark is a set of new deep learning models trained on different tasks, such as privacy feedback classification, privacy issues generation, and high-level theme creation. We illustrate Hark’s efficacy on a corpus of 626 M Google Play reviews. Out of this corpus, our privacy feedback classifier extracts privacy-related reviews (with an AUC-ROC of 0.92). With three annotation studies, we show that Hark’s generated issues are of high accuracy and coverage and that the theme titles are of high quality. We illustrate Hark’s capabilities by presenting high-level insights from Android apps.1an English verb meaning to “pay close attention”
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Stuck in the Permissions With You: Developer & End-User Perspectives on App Permissions & Their Privacy RamificationsMohammad Tahaei, Ruba Abu-Salma, Awais RashidCHI 2023 · 被引用 36 次
- A NEW HOPE: Contextual Privacy Policies for Mobile Applications and An Approach Toward Automated GenerationShidong Pan, Zhen Tao, Thong Hoang, Dawen Zhang 等USENIX Security 2024 · 被引用 24 次
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- A Decade of Privacy-Relevant Android App Reviews: Large Scale TrendsOmer Akgul, Sai Teja Peddinti, Nina Taft, Michelle L. Mazurek 等USENIX Security 2024 · 被引用 14 次
- Shortchanged: Uncovering and Analyzing Intimate Partner Financial Abuse in Consumer ComplaintsArkaprabha Bhattacharya, Kevin Lee, Vineeth Ravi, Jessica Staddon 等CHI 2024 · 被引用 7 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Analyzing User Perspectives on Mobile App Privacy at ScalePreksha Nema, Pauline Anthonysamy, Nina Taft, Sai Teja PeddintiICSE 2022 · 被引用 50 次
- GoEmotions: A Dataset of Fine-Grained EmotionsDorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan S. Cowen 等ACL 2020 · 被引用 16 次
相关 Paper
- Unsupervised Summarization of Privacy Concerns in Mobile Application ReviewsFahimeh Ebrahimi, Anas MahmoudASE 2022 · 被引用 18 次
- Privacy at Scale: Introducing the PrivaSeer Corpus of Web Privacy PoliciesMukund Srinath, Shomir Wilson, C. Lee GilesACL 2021
- Short Text, Large Effect: Measuring the Impact of User Reviews on Android App Security & PrivacyDuc Cuong Nguyen, Erik Derr, Michael Backes, Sven BugielS&P 2019 · 被引用 66 次
- Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep LearningHamza Harkous, Kassem Fawaz, Rémi Lebret, Florian Schaub 等USENIX Security 2018 · 被引用 400 次
- Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated ApproachQinge Xie, Karthik Ramakrishnan, Frank LiUSENIX Security 2025
