Incentivizing Truthfulness Through Audits in Strategic Classification
Andrew Estornell, Sanmay Das, Yevgeniy Vorobeychik
Abstract
In many societal resource allocation domains, machine learning methods are increasingly used to either score or rank agents in order to decide which ones should receive either resources (e.g., homeless services) or scrutiny (e.g., child welfare investigations) from social services agencies. An agency's scoring function typically operates on a feature vector that contains a combination of self-reported features and information available to the agency about individuals or households. This can create incentives for agents to misrepresent their self-reported features in order to receive resources or avoid scrutiny, but agencies may be able to selectively audit agents to verify the veracity of their reports. We study the problem of optimal auditing of agents in such settings. When decisions are made using a threshold on an agent's score, the optimal audit policy has a surprisingly simple structure, uniformly auditing all agents who could benefit from lying. While this policy can, in general be hard to compute because of the difficulty of identifying the set of agents who could benefit from lying given a complete set of reported types, we also present necessary and sufficient conditions under which it is tractable. We show that the scarce resource setting is more difficult, and exhibit an approximately optimal audit policy in this case. In addition, we show that in either setting verifying whether it is possible to incentivize exact truthfulness is hard even to approximate. However, we also exhibit sufficient conditions for solving this problem optimally, and for obtaining good approximations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DeRDaVa: Deletion-Robust Data Valuation for Machine LearningXiao Tian, Rachael Hwee Ling Sim, Jue Fan, Bryan Kian Hsiang LowAAAI 2024 · 3 citations
- On the Computational Complexity of Performative PredictionIoannis Anagnostides, Rohan Chauhan, Ioannis Panageas, Tuomas Sandholm et al.ICML 2026 · 1 citation
- Disentangling misreporting from genuine adaptation in strategic settings: a causal approachDylan Zapzalka, Trenton Chang, Lindsay A. Warrenburg, Sae-Hwan Park et al.NeurIPS 2025 · 1 citation
- Optimally Auditing Adversarial AgentsSanmay Das, Fang-Yi Yu, Yuang ZhangAAAI 2026 · 1 citation
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Automatically Evading Classifiers: A Case Study on PDF Malware ClassifiersWeilin Xu, Yanjun Qi, David EvansNDSS 2016 · 249 citations
- Defending Against Physically Realizable Attacks on Image ClassificationTong Wu, Liang Tong, Yevgeniy VorobeychikICLR 2020 · 143 citations
- Improving Robustness of ML Classifiers against Realizable Evasion Attacks Using Conserved FeaturesLiang Tong, Bo Li, Chen Hajaj, Chaowei Xiao et al.USENIX Security 2019 · 95 citations
Related papers
- Classification with Strategically Withheld DataAnilesh K. Krishnaswamy, Haoming Li, David Rein, Hanrui Zhang et al.AAAI 2021 · 17 citations
- Comparing Targeting Strategies for Maximizing Social Welfare with Limited ResourcesVibhhu Sharma, Bryan WilderICLR 2025
- Designing Optimal Mechanisms to Locate Facilities with Insufficient Capacity for Bayesian AgentsGennaro Auricchio, Jie ZhangAAAI 2026
- Incentive-Aware Dynamic Resource Allocation under Long-Term Cost ConstraintsYan Dai, Negin Golrezaei, Patrick JailletNeurIPS 2025
- Heterogeneous Facility Location with Limited ResourcesArgyrios Deligkas, Aris Filos-Ratsikas, Alexandros A. VoudourisAAAI 2022 · 29 citations
