Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistants
Yash Vekaria, Aurelio Loris Canino, Jonathan Levitsky, Alex Ciechonski, Patricia Callejo, Anna Maria Mandalari, Zubair Shafiq
摘要
Generative AI (GenAI) browser assistants integrate powerful capabilities of GenAI in web browsers to provide rich experiences such as question answering, content summarization, and agentic navigation. These assistants, available today as browser extensions, can not only track detailed browsing activity such as search and click data, but can also autonomously perform tasks such as filling forms, raising significant privacy concerns. It is crucial to understand the design and operation of GenAI browser extensions, including how they collect, store, process, and share user data. To this end, we study their ability to profile users and personalize their responses based on explicit or inferred demographic attributes and interests of users. We perform network traffic analysis and use a novel prompting framework to audit tracking, profiling, and personalization by the ten most popular GenAI browser assistant extensions. We find that instead of relying on local in-browser models, these assistants largely depend on server-side APIs, which can be auto-invoked without explicit user interaction. When invoked, they collect and share webpage content, often the full HTML DOM and sometimes even the user's form inputs, with their first-party servers. Some assistants also share identifiers and user prompts with third-party trackers such as Google Analytics. The collection and sharing continues even if a webpage contains sensitive information such as health or personal information such as name or SSN entered in a web form. We find that several GenAI browser assistants infer demographic attributes such as age, gender, income, and interests and use this profile--which carries across browsing contexts--to personalize responses. In summary, our work shows that GenAI browser assistants can and do collect personal and sensitive information for profiling and personalization with little to no safeguards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Beyond Memorization: Violating Privacy via Inference with Large Language ModelsRobin Staab, Mark Vero, Mislav Balunovic, Martin T. VechevICLR 2024 · 被引用 211 次
- Mystique: Uncovering Information Leakage from Browser ExtensionsQuan Chen, Alexandros KapravelosCCS 2018 · 被引用 88 次
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan 等VLDB 2024 · 被引用 66 次
- DoubleX: Statically Detecting Vulnerable Data Flows in Browser Extensions at ScaleAurore Fass, Dolière Francis Somé, Michael Backes, Ben StockCCS 2021 · 被引用 35 次
- Are they Toeing the Line? Diagnosing Privacy Compliance Violations among Browser ExtensionsYuxi Ling, Kailong Wang, Guangdong Bai, Haoyu Wang 等ASE 2022 · 被引用 17 次
相关 Paper
- A Comparative Study of Users' Information-Seeking Practices Across Search Engines and Generative AI ChatbotsElsa Lichtenegger, Aleksandra Urman, Aniko HannakSIGIR 2026
- Facilitating Proactive and Reactive Guidance for Decision Making on the Web: A Design Probe with WebSeekYanwei Huang, Arpit NarechaniaCHI 2026 · 被引用 2 次
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI OverviewsRiley Grossman, Songjiang Liu, Michael K. Chen, Mike Smith 等SIGIR 2026
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal InformationXiao Zhan, Juan Carlos Carrillo, William Seymour, Jose SuchUSENIX Security 2025
- Understanding the Performance Costs and Benefits of Privacy-focused Browser ExtensionsKevin Borgolte, Nick FeamsterWWW 2020 · 被引用 21 次
