Leveraging Similar Users for Personalized Language Modeling with Limited Data
Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
Abstract
Personalized language models are designed and trained to capture language patterns specific to individual users. This makes them more accurate at predicting what a user will write. However, when a new user joins a platform and not enough text is available, it is harder to build effective personalized language models. We propose a solution for this problem, using a model trained on users that are similar to a new user. In this paper, we explore strategies for finding the similarity between new users and existing ones and methods for using the data from existing users who are a good match. We further explore the trade-off between available data for new users and how well their language can be modeled.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ec2930f-95d4-4049-a4a6-09a95dc7144aCited by top-tier papers7
- Large Language Models Empowered Personalized Web AgentsHongru Cai, Yongqi Li, Wenjie Wang, Fengbin Zhu et al.WWW 2025 · 62 citations
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng et al.ICLR 2024 · 46 citations
- Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory GraphsZheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu et al.EMNLP 2024 · 6 citations
- Unifying Data Perspectivism and Personalization: An Application to Social NormsJoan Plepi, Béla Neuendorf, Lucie Flek, Charles WelchEMNLP 2022 · 4 citations
- From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLPAdithya V. Ganesan, Vasudha Varadarajan, Oscar N. E. Kjell, Whitney Ringwald et al.ACL 2026 · 3 citations
Builds on5
- Learning Architectures from an Extended Search Space for Language ModelingYinqiao Li, Chi Hu, Yuhao Zhang, Nuo Xu et al.ACL 2020 · 12 citations
- Refocusing on Relevance: Personalization in NLGShiran Dudy, Steven Bedrick, Bonnie WebberEMNLP 2021 · 2 citations
- Compositional Demographic Word EmbeddingsCharles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada MihalceaEMNLP 2020
- Selecting Informative Contexts Improves Language Model Fine-tuningRichard J. Antonello, Nicole Beckage, Javier Turek, Alexander HuthACL 2021
- Dynamic Contextualized Word EmbeddingsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
Related papers
- Learning from My Friends: Few-Shot Personalized Conversation Systems via Social NetworksZhiliang Tian, Wei Bi, Zihan Zhang, Dongkyu Lee et al.AAAI 2021 · 12 citations
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni et al.ICML 2025
- LLMs + Persona-Plug = Personalized LLMsJiongnan Liu, Yutao Zhu, Shuting Wang, Xiaochi Wei et al.ACL 2025 · 19 citations
- Future Language Modeling from Temporal Document HistoryChangmao Li, Jeffrey FlaniganICLR 2024
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
