Personalized Language Model Learning on Text Data Without User Identifiers
Yucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu, Fandong Meng, Jie Zhou, Ning Liu, Fan Wu, Guihai Chen
Abstract
In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers. However, the absence of user identifiers restricts the ability of cloud-based language models to provide personalized services, which are essential for catering to diverse user needs. The trivial method of replacing an explicit user identifier with a static user embedding as model input still compromises data anonymization. In this work, we propose to let each mobile device maintain a user-specific distribution to dynamically generate user embeddings, thereby breaking the one-to-one mapping between an embedding and a specific user. We further theoretically demonstrate that to prevent the cloud from tracking users via uploaded embeddings, the local distributions of different users should either be derived from a linearly dependent space to avoid identifiability or be close to each other to prevent accurate attribution. Evaluation on both public and industrial datasets using different language models reveals a remarkable improvement in accuracy from incorporating anonymous user embeddings, while preserving real-time inference requirement. CCS Concepts • Information systems → Data mining; • Computing methodologies → Natural language processing; • Human-centered computing → Ubiquitous and mobile computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a9758df-1403-4fa5-ab44-d787fde802aeCited by top-tier papers1
Ask how each one uses itBuilds on10
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning ApproachAlireza Fallah, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2020 · 1,354 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Exploiting Shared Representations for Personalized Federated LearningLiam Collins, Hamed Hassani, Aryan Mokhtari, Sanjay ShakkottaiICML 2021 · 1,081 citations
- Personalized Federated Learning via Variational Bayesian InferenceXu Zhang, Yinchuan Li, Wenpeng Li, Kaiyang Guo et al.ICML 2022 · 132 citations
Related papers
- TextFusion: Privacy-Preserving Pre-trained Model Inference via Token FusionXin Zhou, Jinzhu Lu, Tao Gui, Ruotian Ma et al.EMNLP 2022 · 12 citations
- Compositional Demographic Word EmbeddingsCharles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada MihalceaEMNLP 2020
- Reconstruction Attack-Resistant Inference Paradigm for LLM Cloud ServicesZipeng Ye, Wenjian Luo, Qi Zhou, Yubo TangAAAI 2026
- Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMsDong Yan, Jian Liang, Ran He, Tieniu TanICLR 2026 · 3 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
