Safety Certificate against Latent Variables with Partially Unidentifiable Dynamics
Haoming Jing, Yorie Nakahira
摘要
Many systems contain latent variables that make their dynamics partially unidentifiable or cause distribution shifts in the observed statistics between offline and online data. However, existing control techniques often assume access to complete dynamics or perfect simulators with fully observable states, which are necessary to verify whether the system remains within a safe set (forward invariance) or safe actions are consistently feasible at all times. To address this limitation, we propose a technique for designing probabilistic safety certificates for systems with latent variables. A key technical enabler is the formulation of invariance conditions in probability space, which can be constructed using observed statistics in the presence of distribution shifts due to latent variables. We use this invariance condition to construct a safety certificate that can be implemented efficiently in real-time control. The proposed safety certificate can continuously find feasible actions that control long-term risk to stay within tolerance. Stochastic safe control and (causal) reinforcement learning have been studied in isolation until now. To the best of our knowledge, the proposed work is the first to use causal reinforcement learning to quantify long-term risk for the design of safety certificates. This integration enables safety certificates to efficiently ensure longterm safety in the presence of latent variables. The effectiveness of the proposed safety certificate is demonstrated in numerical simulations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Provably Efficient Causal Reinforcement Learning with Confounded Observational DataLingxiao Wang, Zhuoran Yang, Zhaoran WangNeurIPS 2021 · 被引用 61 次
- Lyapunov Density Models: Constraining Distribution Shift in Learning-Based ControlKatie Kang, Paula Gradu, Jason J. Choi, Michael Janner 等ICML 2022 · 被引用 39 次
- A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision ProcessesChengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan JiangICML 2022 · 被引用 31 次
- An Instrumental Variable Approach to Confounded Off-Policy EvaluationYang Xu, Jin Zhu, Chengchun Shi, Shikai Luo 等ICML 2023 · 被引用 24 次
- Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric ModelsRui Miao, Zhengling Qi, Xiaoke ZhangNeurIPS 2022 · 被引用 18 次
相关 Paper
- Learning Control Policies for Stochastic Systems with Reach-Avoid GuaranteesDorde Zikelic, Mathias Lechner, Thomas A. Henzinger, Krishnendu ChatterjeeAAAI 2023 · 被引用 50 次
- Stochastic Minimum-Cost Reach-Avoid Reinforcement LearningJingduo Pan, Taoran Wu, Yiling Xue, Bai XueICML 2026
- Infinite Time Horizon Safety of Bayesian Neural NetworksMathias Lechner, Dorde Zikelic, Krishnendu Chatterjee, Thomas A. HenzingerNeurIPS 2021 · 被引用 20 次
- Neural Control and Certificate Repair via Runtime MonitoringEmily Yu, Dorde Zikelic, Thomas A. HenzingerAAAI 2025 · 被引用 5 次
- Quantitative Supermartingale CertificatesAlessandro Abate, Mirco Giacobbe, Diptarko RoyCAV 2025 · 被引用 7 次
