Compositional Demographic Word Embeddings
Charles Welch, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
Abstract
Word embeddings are usually derived from corpora containing text from many individuals, thus leading to general purpose representations rather than individually personalized representations. While personalized embeddings can be useful to improve language model performance and other language processing tasks, they can only be computed for people with a large amount of longitudinal data, which is not the case for new users. We propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user (i.e., gender, age, location, religion). We show that the resulting demographic-aware word representations outperform generic word representations on two tasks for English: language modeling and word associations. We further explore the trade-off between the number of available attributes and their relative effectiveness and discuss the ethical implications of using them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 759c04b7-985f-4b42-83dc-6a11092a670cCited by top-tier papers5
- Leveraging Similar Users for Personalized Language Modeling with Limited DataCharles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas et al.ACL 2022 · 37 citations
- Unifying Data Perspectivism and Personalization: An Application to Social NormsJoan Plepi, Béla Neuendorf, Lucie Flek, Charles WelchEMNLP 2022 · 4 citations
- Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language ModelsYi Zhou, Danushka Bollegala, José Camacho-ColladosEMNLP 2024 · 3 citations
- Learning Dynamic Contextualised Word Embeddings via Template-based Temporal AdaptationXiaohang Tang, Yi Zhou, Danushka BollegalaACL 2023 · 2 citations
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
Builds on1
Related papers
- Employing Personal Word Embeddings for Personalized SearchJing Yao, Zhicheng Dou, Ji-Rong WenSIGIR 2020 · 41 citations
- Dynamic Contextualized Word EmbeddingsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
- Personalized Language Model Learning on Text Data Without User IdentifiersYucheng Ding, Yangwenjian Tan, Xiangyu Liu, Chaoyue Niu et al.KDD 2025 · 2 citations
- LLMs + Persona-Plug = Personalized LLMsJiongnan Liu, Yutao Zhu, Shuting Wang, Xiaochi Wei et al.ACL 2025 · 19 citations
- SocioProbe: What, When, and Where Language Models Learn about SociodemographicsAnne Lauscher, Federico Bianchi, Samuel R. Bowman, Dirk HovyEMNLP 2022 · 6 citations
