Reliability Testing for Natural Language Processing Systems
Samson Tan, Shafiq R. Joty, Kathy Baxter, Araz Taeihagh, Gregory A. Bennett, Min-Yen Kan
Abstract
Questions of fairness, robustness, and transparency are paramount to address before deploying NLP systems. Central to these concerns is the question of reliability: Can NLP systems reliably treat different demographics fairly and function correctly in diverse and noisy environments? To address this, we argue for the need for reliability testing and contextualize it among existing work on improving accountability. We show how adversarial attacks can be reframed for this goal, via a framework for developing reliability tests. We argue that reliability testing -with an emphasis on interdisciplinary collaboration -will enable rigorous and targeted testing, and aid in the enactment and enforcement of industry standards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8eb9f780-7030-4b5c-9490-5bb487d7388bCited by top-tier papers4
- Prompting GPT-3 To Be ReliableChenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang et al.ICLR 2023 · 68 citations
- TAI3: Testing Agent Integrity in Interpreting User IntentShiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang et al.NeurIPS 2025 · 9 citations
- LEAP: Efficient and Automated Test Method for NLP SoftwareMingxuan Xiao, Yan Xiao, Hai Dong, Shunhui Ji et al.ASE 2023 · 7 citations
- On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty AgentsJen-tse Huang, Jiaxu Zhou, Tailin Jin, Xuhui Zhou et al.ICML 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- The King Is Naked: On the Notion of Robustness for Natural Language ProcessingEmanuele La Malfa, Marta KwiatkowskaAAAI 2022 · 31 citations
- Fairness Beyond Performance: Revealing Reliability Disparities Across Groups in Legal NLPT. Y. S. S. Santosh, Irtiza ChowdhuryACL 2025
- ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language ModelsAparna Elangovan, Ling Liu, Lei Xu, Sravan Babu Bodapati et al.ACL 2024 · 19 citations
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted AttackBoxin Wang, Hengzhi Pei, Boyuan Pan, Qian Chen et al.EMNLP 2020 · 55 citations
- The Impossibility of Fair LLMsJacy Reese Anthis, Kristian Lum, Michael D. Ekstrand, Avi Feller et al.ACL 2025
