Perturbation Effects on Robustness and Individual Fairness
Xuran Li, Hao Xue, Peng Wu, Xingjun Ma, Zhen Zhang, Huaming Chen, Flora D. Salim
Abstract
Deep neural networks are vulnerable to adversarial perturbations that can simultaneously degrade prediction robustness and individual fairness across diverse application settings. However, existing evaluation protocols typically assess these dimensions in isolation, thereby obscuring critical failure modes. To bridge this gap, we formalize Robust Individual Fairness (RIF): under semantic-preserving (truth-condition-preserving) perturbations, predictions should remain both correct with respect to the ground truth and invariant across semantically equivalent individuals. To surface RIF violations in practice, we introduce RIFair, a black-box adversarial framework that leverages a decoupled perturbation strategy to construct semantically preserved yet unrobust and/or unfair instance pairs. Experiments across multiple model architectures and real-world textual datasets show that robustness-only or fairness-only metrics often miss Robust Biased and Unrobust Fair behaviors. RIFair reliably exposes these hidden vulnerabilities, supporting RIF as a necessary criterion for trustworthy model assessment. The experimental code is publicly available at https://github.com/Xuran-LI/RIFair.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 888f656b-cc1e-467a-900e-805b4bce3ac0Builds on9
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong et al.ICSE 2020 · 127 citations
- Efficient white-box fairness testing through gradient searchLingfeng Zhang, Yueling Zhang, Min ZhangISSTA 2021 · 51 citations
Related papers
- Is Fairness Truly Fair? Towards Reliable Lipschitz Fairness in Multi-Task Learning via Fixed-δ AlignmentJunbo Ding, Xin Zang, Chenchen Pan, Donghao Song et al.KDD 2026
- Verifying Neural Network Robustness with Dual PerturbationsHai Duong, Lam Nguyen, Thanh Le, ThanhVu NguyenCVPR 2026 · 4 citations
- Proactive Defense Benchmark against Deepfake GenerationJoonhyuk Baek, Wonjune Seo, Jae-yun Kim, Saerom Park et al.ICML 2026
- MAFT: Efficient Model-Agnostic Fairness Testing for Deep Neural Networks via Zero-Order Gradient SearchZhaohui Wang, Min Zhang, Jingran Yang, Bojie Shao et al.ICSE 2024 · 6 citations
- Your Neighbor Matters: Towards Fair Decisions Under Networked InterferenceWenjing Yang, Haotian Wang, Haoxuan Li, Hao Zou et al.KDD 2024 · 2 citations
