Closing the Sim-to-Real Gap in Non-Markovian Spreading Processes via GPU-Accelerated Distributional RL
Heman Shakeri
摘要
Controlling spreading processes on networks, such as epidemics, information cascades, and product adoption, requires policies that perform on realistic stochastic dynamics, not just tractable approximations. Yet policies trained on standard simplifications (mean-field ODEs, Markovian dynamics) suffer severe performance degradation at deployment. We trace this sim-to-real gap to three theoretical pathologies: Optimism Bias, where deterministic approximations systematically underestimate variance via Jensen's inequality; Hub Blindness, where global state aggregation obscures the super-spreaders driving scale-free networks; and the Valley of Death, where mean-value critics fail to navigate the bimodal nature (extinction vs. viral) of cascade outcomes. We resolve these challenges through two synergistic contributions. First, the Stratified Mean-Field Observer partitions nodes by influence tier, preserving hub dynamics at cost while producing fixed-dimensional observations that enable zero-shot transfer across network scales and topologies. Second, we show that distributional RL via Truncated Quantile Critics improves risk-aware control of bimodal cascades. Trained on a GPU-accelerated simulator supporting non-Markovian renewal dynamics, our approach achieves improvement over Markovian baselines and robust zero-shot transfer to real-world social networks (Facebook, Twitter, YouTube), significantly mitigating the simulation-to-reality gap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- Controlling Graph Dynamics with Reinforcement Learning and Graph Neural NetworksEli A. Meirom, Haggai Maron, Shie Mannor, Gal ChechikICML 2021 · 被引用 56 次
- A Benchmark Study of Deep-RL Methods for Maximum Coverage Problems over GraphsZhicheng Liang, Yu Yang, Xiangyu Ke, Xiaokui Xiao 等VLDB 2024 · 被引用 3 次
相关 Paper
- Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes TestbedMinjae Kwon, Josephine Lamp, Lu FengICML 2026 · 被引用 1 次
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
- Network Diffusions via Neural Mean-Field DynamicsShushan He, Hongyuan Zha, Xiaojing YeNeurIPS 2020 · 被引用 10 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RLAndrew Wagenmaker, Kevin Huang, Liyiming Ke, Kevin Jamieson 等NeurIPS 2024 · 被引用 45 次
