Robust Estimation of Population-Level Effects in Repeated-Measures NLP Experimental Designs
Alejandro Benito-Santos, Adrián Ghajari, Víctor Fresno
Abstract
NLP research frequently grapples with multiple sources of variability-spanning runs, datasets, annotators, and more-yet conventional analysis methods often neglect these hierarchical structures, threatening the reproducibility of findings. To address this gap, we contribute a case study illustrating how linear mixed-effects models (LMMs) can rigorously capture systematic language-dependent differences (i.e., population-level effects) in a population of monolingual and multilingual language models. In the context of a bilingual hate speech detection task, we demonstrate that LMMs can uncover significant population-level effects-even under low-resource (small-N) experimental designs-while mitigating confounds and random noise. By setting out a transparent blueprint for repeated-measures experimentation, we encourage the NLP community to embrace variability as a feature, rather than a nuisance, in order to advance more robust, reproducible, and ultimately trustworthy results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed80ee8b-47b3-4f57-b2dd-a0ab2fa52c54Builds on5
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia et al.EMNLP 2020 · 76 citations
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang et al.EMNLP 2023 · 2 citations
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Evaluating Extreme Hierarchical Multi-label ClassificationEnrique Amigó, Agustín D. DelgadoACL 2022
- Translating away Translationese without Parallel DataRricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van GenabithEMNLP 2023
Related papers
- Hate Personified: Investigating the role of LLMs in content moderationSarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser et al.EMNLP 2024 · 6 citations
- Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine TrabelsiEMNLP 2025 · 1 citation
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 16 citations
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke et al.ACL 2023 · 23 citations
- LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and TargetMd. Arid Hasan, Firoj Alam, Md Fahad Hossain, Usman Naseem et al.ACL 2026
