Lune

ACL2020Top-tier venue

Beyond Accuracy: Behavioral Testing of NLP Models with CheckList

Marco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer Singh

2020Year
51Citations
225Top-tier citations

Abstract

Although measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors. Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a taskagnostic methodology for testing NLP models. CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly. We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models. In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model. In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b03a51d3-c999-4198-b6f1-6423e472cc3f

Cited by top-tier papers225

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines