Lune

SOSP2026Top-tier venue

A nchor : Mitigating GPU Shallow Disruptions with Decoupled Memory

Haoyi Ma, Shiwei Gao, Youmin Chen, Junrong Huang, Youyou Lu, Jiwu Shu

2026Year

Abstract

Job failures are frequent in large-scale GPU clusters for LLM workloads, leading to significant resource wastage. The vast majority of these are shallow disruptions (e.g., software errors or updates), where only the worker process crashes while the underlying GPU and OS kernel remain intact. Existing recovery systems, however, are failure-agnostic, relying on heavyweight checkpointing that forces a full, slow state reload even when the data could have survived on the GPU.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get e3725449-6e8c-4842-b4e1-c1bbc58bc16d

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines