Robust Estimation of Population-Level Effects in Repeated-Measures NLP Experimental Designs
Alejandro Benito-Santos, Adrián Ghajari, Víctor Fresno
摘要
NLP research frequently grapples with multiple sources of variability-spanning runs, datasets, annotators, and more-yet conventional analysis methods often neglect these hierarchical structures, threatening the reproducibility of findings. To address this gap, we contribute a case study illustrating how linear mixed-effects models (LMMs) can rigorously capture systematic language-dependent differences (i.e., population-level effects) in a population of monolingual and multilingual language models. In the context of a bilingual hate speech detection task, we demonstrate that LMMs can uncover significant population-level effects-even under low-resource (small-N) experimental designs-while mitigating confounds and random noise. By setting out a transparent blueprint for repeated-measures experimentation, we encourage the NLP community to embrace variability as a feature, rather than a nuisance, in order to advance more robust, reproducible, and ultimately trustworthy results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- With Little Power Comes Great ResponsibilityDallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia 等EMNLP 2020 · 被引用 76 次
- We Need to Talk About Reproducibility in NLP Model ComparisonYan Xue, Xuefei Cao, Xingli Yang, Yu Wang 等EMNLP 2023 · 被引用 2 次
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Evaluating Extreme Hierarchical Multi-label ClassificationEnrique Amigó, Agustín D. DelgadoACL 2022
- Translating away Translationese without Parallel DataRricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van GenabithEMNLP 2023
相关 Paper
- Hate Personified: Investigating the role of LLMs in content moderationSarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser 等EMNLP 2024 · 被引用 6 次
- Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine TrabelsiEMNLP 2025 · 被引用 1 次
- Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced LanguagesPaul Röttger, Debora Nozza, Federico Bianchi, Dirk HovyEMNLP 2022 · 被引用 16 次
- NLPositionality: Characterizing Design Biases of Datasets and ModelsSebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke 等ACL 2023 · 被引用 23 次
- LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and TargetMd. Arid Hasan, Firoj Alam, Md Fahad Hossain, Usman Naseem 等ACL 2026
