DynaSent: A Dynamic Benchmark for Sentiment Analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, Douwe Kiela
Abstract
We introduce DynaSent ('Dynamic Sentiment'), a new English-language benchmark task for ternary (positive/negative/neutral) sentiment analysis. DynaSent combines naturally occurring sentences with sentences created using the open-source Dynabench Platform, which facilities human-and-model-inthe-loop dataset creation. DynaSent has a total of 121,634 sentences, each validated by five crowdworkers, and its development and test splits are designed to produce chance performance for even the best models we have been able to develop; when future models solve this task, we will use them to create DynaSent version 2, continuing the dynamic evolution of this benchmark. Here, we report on the dataset creation effort, focusing on the steps we took to increase quality and reduce artifacts. We also present evidence that DynaSent's Neutral category is more coherent than the comparable category in other benchmarks, and we motivate training models from scratch for each round over successive fine-tuning. * Equal contribution. Model 0 RoBERTa fine-tuned on sentiment benchmarks Model 0 used to find challenging naturally occurring sentences Human validation Round 1 Dataset Model 1 RoBERTa fine-tuned on sentiment benchmarks + Round 1 Dataset Dynabench used to crowdsource sentences that fool Model 1 Human validation Round 2 Dataset
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal et al.NeurIPS 2021 · 315 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- Human-Adversarial Visual Question AnsweringSasha Sheng, Amanpreet Singh, Vedanuj Goswami, Jose Alberto Lopez Magana et al.NeurIPS 2021 · 81 citations
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain et al.NeurIPS 2021 · 76 citations
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
Builds on2
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Leveraging Affirmative Interpretations from Negation Improves Natural Language UnderstandingMd Mosharaf Hossain, Eduardo BlancoEMNLP 2022 · 4 citations
- Rather a Nurse than a Physician - Contrastive Explanations under InvestigationOliver Eberle, Ilias Chalkidis, Laura Cabello, Stephanie BrandlEMNLP 2023 · 3 citations
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
- SentiBERT: A Transferable Transformer-Based Architecture for Compositional Sentiment SemanticsDa Yin, Tao Meng, Kai-Wei ChangACL 2020 · 127 citations
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
