Accurate and Ultra-Fast Launch-Time Validation of Idempotency for GPU Kernels
Mingcong Han, Weihang Shen, Rong Chen, Haibo Chen
Abstract
We discovered that a GPU kernel can have both idempotent and non-idempotent instances depending on the input. These kernels, called conditionally-idempotent, are common in real-world GPU applications—490 out of 547 from six popular applications. This finding reveals a limitation in previous work that statically classifies GPU kernels as idempotent or non-idempotent, potentially compromising the correctness and effectiveness of idempotence-based systems. This paper presents Picker, the first launch-time analysis system for instance-level idempotency validation. Picker accurately validates the idempotency of GPU kernel instances before execution by utilizing launch arguments. Several optimizations are proposed to reduce validation latency to microseconds. Evaluations using representative GPU applications (547 kernels and 18,217 instances) show that Picker accurately identifies idempotent instances with zero false positives and an 18.54% false-negative rate. The launch-time validation completes in under 5 μs for all instances (about 90% under 1 μs). Through integration, Picker reduces checkpoint costs to less than 4% in fault-tolerant systems and decreases preemption latency by 84.2% in scheduling systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 266f2c34-a5f3-4a65-941c-72db6bd41131Related papers
- Featherweight Soft Error Resilience for GPUsYida Zhang, Changhee JungMICRO 2022 · 14 citations
- sfGPUMC: A Stateless Model Checker for GPU Weak Memory ConcurrencySoham Chakraborty, S. Krishna, Andreas Pavlogiannis, Omkar TuppeCAV 2025 · 2 citations
- Microsecond-scale Preemption for Concurrent GPU-accelerated DNN InferencesMingcong Han, Hanze Zhang, Rong Chen, Haibo ChenOSDI 2022 · 153 citations
- A Modular Static Cost Analysis for GPU Warp-Level ParallelismGregory Blike, Hannah Zicarelli, Udaya Sathiyamoorthy, Julien Lange et al.POPL 2026 · 1 citation
- Equivalence Checking of ML GPU KernelsBenjamin Driscoll, Kshitij Dubey, Anjiang Wei, Neeraj Kayal et al.OOPSLA 2026 · 1 citation
