Closing the Sim-to-Real Gap in Non-Markovian Spreading Processes via GPU-Accelerated Distributional RL
Heman Shakeri
Abstract
Controlling spreading processes on networks, such as epidemics, information cascades, and product adoption, requires policies that perform on realistic stochastic dynamics, not just tractable approximations. Yet policies trained on standard simplifications (mean-field ODEs, Markovian dynamics) suffer severe performance degradation at deployment. We trace this sim-to-real gap to three theoretical pathologies: Optimism Bias, where deterministic approximations systematically underestimate variance via Jensen's inequality; Hub Blindness, where global state aggregation obscures the super-spreaders driving scale-free networks; and the Valley of Death, where mean-value critics fail to navigate the bimodal nature (extinction vs. viral) of cascade outcomes. We resolve these challenges through two synergistic contributions. First, the Stratified Mean-Field Observer partitions nodes by influence tier, preserving hub dynamics at cost while producing fixed-dimensional observations that enable zero-shot transfer across network scales and topologies. Second, we show that distributional RL via Truncated Quantile Critics improves risk-aware control of bimodal cascades. Trained on a GPU-accelerated simulator supporting non-Markovian renewal dynamics, our approach achieves improvement over Markovian baselines and robust zero-shot transfer to real-world social networks (Facebook, Twitter, YouTube), significantly mitigating the simulation-to-reality gap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3c0998e-0136-4709-abac-4cceb6c71c94Builds on3
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Controlling Graph Dynamics with Reinforcement Learning and Graph Neural NetworksEli A. Meirom, Haggai Maron, Shie Mannor, Gal ChechikICML 2021 · 56 citations
- A Benchmark Study of Deep-RL Methods for Maximum Coverage Problems over GraphsZhicheng Liang, Yu Yang, Xiangyu Ke, Xiaokui Xiao et al.VLDB 2024 · 3 citations
Related papers
- Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes TestbedMinjae Kwon, Josephine Lamp, Lu FengICML 2026 · 1 citation
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
- Network Diffusions via Neural Mean-Field DynamicsShushan He, Hongyuan Zha, Xiaojing YeNeurIPS 2020 · 10 citations
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 86 citations
- Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RLAndrew Wagenmaker, Kevin Huang, Liyiming Ke, Kevin Jamieson et al.NeurIPS 2024 · 45 citations
