Lune

KDD2026顶会

Continuous-Time Counterfactual Quantile Learning for Risk-Sensitive Policy Optimization

Yi He, Anpeng Wu, Ruoxuan Xiong, Yingrong Wang, Kun Kuang

2026年份

摘要

This paper studies the problem of Continuous-Time Counterfactual Quantile Learning (CT-CQL) for risk-sensitive policy optimization. In many real-world applications such as patient blood pressure monitoring, financial market analysis, and autonomous driving, data is high-frequency and continuously evolving. However, most existing causal inference methods focus on expectation-based or discrete-time counterfactual reasoning, which fail to capture fine-grained temporal dynamics. As a result, policies optimized under these frameworks may overlook critical risks—e.g., a treatment policy with good average outcomes may still expose patients to life-threatening episodes. To overcome these limitations, we propose CT-CQL, a framework built upon a novel identification theory and featuring three key components: (1) modeling full counterfactual outcome distributions via Stochastic Differential Equations (SDEs) governed by the Fokker–Planck Equation (FPE); (2) enhancing robustness through a minimax objective that minimizes FPE residuals under adversarial perturbations; and (3) mitigating confounding bias using a double-robust AIPW loss. CT-CQL enables robust policy optimization by identifying optimal intervention strategies that maximize expected utility while adhering to real-world budget and safety constraints. Experiments on widely-used benchmarks and the real-world MIMIC-III dataset demonstrate the validation and superiority of the proposed method. The project is available at: https://github.com/Eliza-YiHe/CT-CQL/

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖