Splits! Flexible Sociocultural Linguistic Investigation at Scale
Eylon Caplan, Tania Chakraborty, Dan Goldwasser
Abstract
Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss healthy eating with words like timing, regularity, and digestion, whereas Americans use vocabulary like balancing food groups and avoiding fat and sugar, reflecting distinct cultural models of nutrition (Banna et al., 2016) . The computational study of these Sociocultural Linguistic Phenomena (SLP) has traditionally been done in NLP via tailored analyses of specific groups or topics, requiring specialized data collection and experimental operationalization-a process not well-suited to quick hypothesis exploration and prototyping. To address this, we propose constructing a "sandbox" designed for systematic and flexible sociolinguistic research. Using our method, we construct a demographically/topically split Reddit dataset, SPLITS!, validated by self-identification and by replicating several known SLPs from existing literature. We showcase the sandbox's utility with a scalable, two-stage process that filters large collections of potential SLPs (PSLPs) to surface the most promising candidates for deeper, qualitative investigation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- VALUE: Understanding Dialect Disparity in NLUCaleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson et al.ACL 2022 · 57 citations
- Studying Politeness across Cultures using English Twitter and Mandarin WeiboMingyang Li, Louis Hickman, Louis Tay, Lyle H. Ungar et al.CSCW 2020 · 43 citations
- Interpretation of NLP models through input marginalizationSiwon Kim, Jihun Yi, Eunji Kim, Sungroh YoonEMNLP 2020 · 41 citations
Related papers
- Deliberate Exposure to Opposing Views and Its Association with Behavior and Rewards on Political CommunitiesAlexandros EfstratiouWWW 2024 · 2 citations
- Culture is Not Trivia: Sociocultural Theory for Cultural NLPNaitian Zhou, David Bamman, Isaac L. BleamanACL 2025 · 33 citations
- Using Sociolinguistic Variables to Reveal Changing Attitudes Towards Sexuality and GenderSky CH-Wang, David JurgensEMNLP 2021 · 6 citations
- Learning Subjective Label Distributions via Sociocultural DescriptorsMohammed Fayiz Parappan, Ricardo HenaoEMNLP 2025 · 5 citations
- MultiPICo: Multilingual Perspectivist Irony CorpusSilvia Casola, Simona Frenda, Soda Marem Lo, Erhan Sezerer et al.ACL 2024 · 2 citations
