Solving Online Threat Screening Games using Constrained Action Space Reinforcement Learning
Sanket Shah, Arunesh Sinha, Pradeep Varakantham, Andrew Perrault, Milind Tambe
Abstract
Large-scale screening for potential threats with limited resources and capacity for screening is a problem of interest at airports, seaports, and other ports of entry. Adversaries can observe screening procedures and arrive at a time when there will be gaps in screening due to limited resource capacities. To capture this game between ports and adversaries, this problem has been previously represented as a Stackelberg game, referred to as a Threat Screening Game (TSG). Given the significant complexity associated with solving TSGs and uncertainty in arrivals of customers, existing work has assumed that screenees arrive and are allocated security resources at the beginning of the time-window. In practice, screenees such as airport passengers arrive in bursts correlated with flight time and are not bound by fixed time-windows. To address this, we propose an online threat screening model in which the screening strategy is determined adaptively as a passenger arrives while satisfying a hard bound on acceptable risk of not screening a threat. To solve the online problem, we first reformulate it as a Markov Decision Process (MDP) in which the hard bound on risk translates to a constraint on the action space and then solve the resultant MDP using Deep Reinforcement Learning (DRL). To this end, we provide a novel way to efficiently enforce linear inequality constraints on the action output in DRL. We show that our solution allows us to significantly reduce screenee wait time without compromising on the risk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement LearningChangyu Chen, Ramesha Karunasena, Thanh Hong Nguyen, Arunesh Sinha et al.NeurIPS 2023 · 16 citations
- Reduced Policy Optimization for Continuous Control with Hard ConstraintsShutong Ding, Jingya Wang, Yali Du, Ye ShiNeurIPS 2023 · 10 citations
- Autoregressive Policy Optimization for Constrained Allocation TasksDavid Winkel, Niklas Strauß, Maximilian Bernhard, Zongyue Li et al.NeurIPS 2024 · 2 citations
- Improving Stochastic Action-Constrained Reinforcement Learning via Truncated DistributionsRoland Stolz, Michael Eichelbeck, Matthias AlthoffAAAI 2026 · 1 citation
Related papers
- Online 3D Bin Packing with Constrained Deep Reinforcement LearningHang Zhao, Qijin She, Chenyang Zhu, Yin Yang et al.AAAI 2021 · 162 citations
- Learning Efficient Online 3D Bin Packing on Packing Configuration TreesHang Zhao, Yang Yu, Kai XuICLR 2022 · 56 citations
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
- Safe Offline Reinforcement Learning with Real-Time Budget ConstraintsQian Lin, Bo Tang, Zifan Wu, Chao Yu et al.ICML 2023 · 31 citations
- DRMD: Deep Reinforcement Learning for Malware Detection Under Concept DriftShae McFadden, Myles Foley, Mario D'Onghia, Chris Hicks et al.AAAI 2026 · 7 citations
