Avoiding Undesired Future with Minimal Cost in Non-Stationary Environments
Wen-Bo Du, Tian Qin, Tian-Zuo Wang, Zhi-Hua Zhou
摘要
Machine learning (ML) has achieved remarkable success in prediction tasks. In many real-world scenarios, rather than solely predicting an outcome using an ML model, the crucial concern is how to make decisions to prevent the occurrence of undesired outcomes, known as the avoiding undesired future (AUF) problem. To this end, a new framework called rehearsal learning has been proposed recently, which works effectively in stationary environments by leveraging the influence relations among variables. In real tasks, however, the environments are usually non-stationary, where the influence relations may be dynamic, leading to the failure of AUF by the existing method. In this paper, we introduce a novel sequential methodology that effectively updates the estimates of dynamic influence relations, which are crucial for rehearsal learning to prevent undesired outcomes in nonstationary environments. Meanwhile, we take the cost of decision actions into account and provide the formulation of AUF problem with minimal action cost under non-stationarity. We prove that in linear Gaussian cases, the problem can be transformed into the well-studied convex quadratically constrained quadratic program (QCQP). In this way, we establish the first polynomial-time rehearsal-based approach for addressing the AUF problem. Theoretical and experimental results validate the effectiveness and efficiency of our method under certain circumstances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Structural Causal Bandits under Markov EquivalenceMin Woo Park, Andy Arditi, Elias Bareinboim, Sanghack LeeNeurIPS 2025 · 被引用 3 次
- Gradient-Based Nonlinear Rehearsal Learning with Multivariate AlterationsTian Qin, Tian-Zuo Wang, Zhi-Hua ZhouAAAI 2025 · 被引用 3 次
- Counterfactual Structural Causal BanditsMin Woo Park, Sanghack LeeICLR 2026 · 被引用 1 次
- Variance-Reduced Long-Term Rehearsal Learning with Quadratic Programming ReformulationWen-Bo Du, Tian Qin, Tian-Zuo Wang, Zhi-Hua ZhouNeurIPS 2025 · 被引用 1 次
- Enabling Optimal Decisions in Rehearsal Learning under CARE ConditionWen-Bo Du, Hao-Yi Lei, Lue Tao, Tian-Zuo Wang 等ICML 2025
它引用的顶会 Paper16
- Identifiability Guarantees for Causal Disentanglement from Soft InterventionsJiaqi Zhang, Kristjan H. Greenewald, Chandler Squires, Akash Srivastava 等NeurIPS 2023 · 被引用 120 次
- Learning Linear Causal Representations from Interventions under General Nonlinear MixingSimon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam 等NeurIPS 2023 · 被引用 113 次
- Linear Causal Disentanglement via InterventionsChandler Squires, Anna Seigal, Salil S. Bhate, Caroline UhlerICML 2023 · 被引用 90 次
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang 等ICML 2022 · 被引用 78 次
- Causal Bandits with Unknown Graph StructureYangyi Lu, Amirhossein Meisami, Ambuj TewariNeurIPS 2021 · 被引用 53 次
相关 Paper
- Rehearsal Learning for Avoiding Undesired FutureTian Qin, Tian-Zuo Wang, Zhi-Hua ZhouNeurIPS 2023 · 被引用 8 次
- Regularizing Second-Order Influences for Continual LearningZhicheng Sun, Yadong Mu, Gang HuaCVPR 2023
- On Measuring Influence in Avoiding Undesired FutureLue Tao, Tian-Zuo Wang, Yuan Jiang, Zhi-Hua ZhouICLR 2026
- Rehearsal revealed: The limits and merits of revisiting samples in continual learningEli Verwimp, Matthias De Lange, Tinne TuytelaarsICCV 2021 · 被引用 121 次
- Pausing Policy Learning in Non-stationary Reinforcement LearningHyunin Lee, Ming Jin, Javad Lavaei, Somayeh SojoudiICML 2024 · 被引用 4 次
