Article

Production Spot GPU LLM Serving: Mitigating Preemption with Interruption Notices, Drain Budgets, and Warm Replacements

Learn how to run resilient, cost-effective LLM serving on Spot GPUs using interruption notice detection, drain budgets, warm replacement, and retry contracts.

Why Are LLMs More Vulnerable to Spot Preemption Than Standard Services?

Spot and preemptible GPU instances typically deliver 60% to 80% cost savings compared to On-Demand instances, making them the primary lever for reducing LLM inference infrastructure bills. However, directly applying generic stateless web app Spot practices to LLM serving frameworks (such as vLLM, TGI, or TensorRT-LLM) often triggers catastrophic availability drops:

  1. High Execution Time: The prefill phase for large prompts and the decoding loop for hundreds to thousands of tokens can run from tens of seconds to several minutes. When an interruption signal arrives, nodes almost always hold deep backlogs of long-running, half-finished requests.
  2. In-Memory KV Cache Residency: Self-attention KV caches reside entirely within VRAM. A hard termination implies that all intermediate computational state is lost; cross-node zero-downtime checkpoints are practically nonexistent.
  3. Irreversible Streaming Tokens: When using Server-Sent Events (SSE) or WebSockets, earlier tokens have already been delivered to the client. If an interrupted request is restarted from scratch on a new node, it readily leads to duplicate responses, branching outputs, or semantic drift. If tool use or function calling is involved, it can trigger duplicate external side effects.

Spot GPU management should not be built on the hope that preemptions won’t happen. Instead, you need a deterministic control pipeline: Signal Ingestion → Admission Close → Drain Budget → Warm Replacement → Idempotent Retry.


Principle 1: Converting the Notice Window into a Drain Budget

Cloud providers offer varying early termination notices, all provided strictly on a best-effort basis:

Cloud / PlatformSignal TypeTypical WindowBehavioral Profile
AWS EC2 SpotInterruption Notice / Rebalance Recommendation~120s (Notice) / Earlier (Recommendation)Rebalance signals carry no SLA; instances are hard-terminated at 120s
Google Cloud (GCP)Metadata Preemption Notice~30s (Default) / Configurable (Preview)Sends an ACPI shutdown signal with an extremely tight default window
Kubernetes (Karpenter)Node Interruption EventDerived from underlying cloud eventsApplies taints automatically and initiates a graceful node drain

In production, you cannot assume you have the entire warning duration. You must calculate the actual Drain Budget by subtracting end-to-end control-plane latency:

$$\text{DrainBudget} = \text{NoticeWindow} - \text{DetectionLatency} - \text{IngressRemovalLatency} - \text{SafetyMargin}$$

  • DetectionLatency: Delay from DaemonSet metadata polling or cloud EventBridge notifications (keep this under 1–3 seconds).
  • IngressRemovalLatency: Time required for EndpointSlice updates, Kube-Proxy syncs, and load balancer/ingress deregistration (typically 2–8 seconds).
  • SafetyMargin: Buffer for network jitter, context flushing, and SIGKILL boundaries (recommend reserving 10–15 seconds).

On AWS with a 120-second window, your actionable Drain Budget is usually 85–95 seconds. Under GCP’s default 30-second window, your realistic budget drops below 15 seconds.


Principle 2: Parallelize Warm Replacement and Draining — Stop Waiting Sequentially

A common anti-pattern in cluster operations is running lifecycle steps sequentially: receive termination signal → drain old Pod → wait for Pod exit → trigger autoscaler scale-out → pull container image and weights → register new Pod. With LLMs, cold starts frequently take several minutes; running these phases serially causes an immediate collapse of cluster serving capacity.

The production-ready architecture splits execution into two parallel paths immediately upon signal confirmation:

                                  ┌── 1. Close admission + race inflight requests within Drain Budget (Old Node)
[Interruption Signal / Taint] ────┤
                                  └── 2. Request replacement instance + warm replacement in parallel (New Node)

By leveraging Karpenter interruption handling or custom autoscaler controllers, you can schedule replacement capacity the instant an evicted node receives a cloud.google.com/gke-preemptible=true or karpenter.sh/disruption=interrupted taint. This maximizes the overlap between bringing the new node up and draining the old node.


Principle 3: Classify Inflight Requests to Cut Losses Early

Once an instance enters the Draining state, it should never passively wait for all requests to finish. Instead, use an estimation algorithm based on remaining tokens and current throughput to decide whether an inflight request can finish within the allocated budget:

def evaluate_inflight_request(req, drain_budget, recent_tpot, safety_factor=0.85):
    """
    Decide whether to finish processing the request locally based on estimated time.
    """
    estimated_remaining_tokens = req.max_tokens - req.generated_tokens
    estimated_remaining_time = estimated_remaining_tokens * recent_tpot
    
    # Continue only if execution fits safely within the discounted Drain Budget
    if estimated_remaining_time <= (drain_budget * safety_factor):
        return "CONTINUE_DRAIN"
    else:
        return "ABORT_AND_FAILOVER"

Route inflight requests based on their operational profiles:

Request ProfileEvaluation CriteriaMitigation Action
Short outputs / Late decodingEst. completion time $\le$ Remaining Drain BudgetContinue decoding and complete before eviction
Long outputs / Early prefillEst. completion time $>$ Remaining Drain BudgetAbort immediately to prevent wasted compute
Read-only idempotent inferenceClient supports retries or gateway manages stateIntercept and reroute to a healthy warm node
Composite / Side-effectingWrites to DB, processes payments, or runs Tool CallsRecord state in gateway ledger; block auto-replay and return explicit status

Principle 4: Streaming Retry Contracts and Gateway Ledgers

Non-streaming JSON requests can be retried simply by verifying idempotency keys. Streaming responses, however, present a major UX challenge: clients have often already received and displayed initial tokens. Starting over from the raw prompt on a new node causes UI flickering, duplicated sentences, and broken experiences.

We recommend deploying a context-aware request ledger inside the API Gateway (e.g., via Envoy, Kong, or a custom ingress proxy):

{
  "request_id": "req_llm_9837421",
  "attempt": 2,
  "served_prefix_tokens": 214,
  "model_revision": "qwen2.5-72b-instruct@v1",
  "prompt_fingerprint": "sha256_e3b0c44...",
  "retry_safe": true,
  "interruption_reason": "spot_eviction"
}

Retry Protocol Strategies:

  1. Clean Break Protocol (Client-Managed): The gateway emits a structured SSE error event: event: error, data: {"code": "SPOT_PREEMPTION", "recoverable": true}. The client-side application captures this event, alerts the user, and automatically sends an append-mode request with the previously rendered text acting as the conversational prefix.
  2. Server-side Prefix Continuation: If the serving engine supports prefix caching and the prompt uses deterministic decoding parameters (temperature=0, fixed seed), the gateway can inject the previously delivered tokens as a prefix into the new replica, streaming back only the delta. This requires strict model version and sampling consistency.

Architecture in Practice: End-to-End Spot Preemption State Machine

+-------------+    Interruption Warning    +--------------+
|    READY    | ------------------------> |   DRAINING   |
+-------------+                           +--------------+
       ^                                    |          |
       |                                    |          | (Parallel Actions)
       |                                    v          v
+-------------+   Model Gate Passed   [Stop Ingress] [Provision Replacement]
|  RECOVERED  | <------------------         |          |
+-------------+                       [Run or Abort]   [Warm Weights & GPU]
                                            |          |
                                            +----+-----+ 
                                                 |
                                                 v
                                          +--------------+
                                          |  TERMINATING |
                                          +--------------+

Critical State Transitions:

  1. Signal Interception (Node-Problem-Detector / Spot Handler): Continuously poll instance metadata. The moment a preemption event is discovered, apply a NoSchedule taint to the Kubernetes node and flip the serving Pod’s readiness status to false.
  2. Application Drain Loop:
async def handle_spot_eviction(deadline_ts):
    serving_engine.stop_accepting_new_requests()
    
    while get_current_time() < (deadline_ts - SAFETY_MARGIN_SECONDS):
        if serving_engine.get_active_request_count() == 0:
            break
        await asyncio.sleep(0.5)
        
    # Abort requests that cannot finish within the budget
    await serving_engine.abort_all_inflight(
        reason="drain_budget_exhausted_spot_cleanup"
    )
  1. Strict Replacement Readiness Gates: Before routing traffic to the replacement instance, never rely solely on standard PodScheduled or ContainersReady conditions. Enforce an end-to-end model verification probe:
    • Validate GPU driver status and memory allocation.
    • Confirm weights are deserialized and bound to device VRAM.
    • Run a dummy prompt execution to warm up CUDA kernels and execution graphs.

Compute Topology: Hybrid Pools and Failure Domain Isolation

Never rely exclusively on a single GPU instance type within a single Availability Zone for your Spot inference workloads. A regional supply squeeze can wipe out your entire pool simultaneously, leading to an availability cascade.

                    ┌── On-Demand Baseline Capacity (Sustains 20%–40% Core Load)
                    │
Inference Cluster ──┼── Spot Pool A (Primary AZ, e.g., A100-80G / AZ-a)
                    ├── Spot Pool B (Alternative Instance/AZ, e.g., L40S / AZ-b)
                    └── Fallback Degradation Pipeline (Load shedding / Quantized model fallback)
  • Baseline Floor: Keep an On-Demand compute baseline dedicated to your minimum viable SLO. If all Spot instances are revoked simultaneously, your service degrades gracefully rather than suffering an outright outage.
  • Instance Diversification: Broaden Karpenter node pool requirements. Allow the scheduler to pick across equivalent GPU families (e.g., fall back across A100, H100, L40S, or A10G depending on model compatibility).

Golden Metrics and Chaos Engineering Checklist

Core Prometheus Metrics

  • llm_spot_notice_to_draining_seconds: Latency between cloud provider interruption emission and the application closing ingress admission.
  • llm_spot_inflight_aborted_ratio: Percentage of requests forcefully terminated due to exhausted drain budgets.
  • llm_spot_replacement_ready_latency_seconds: Wall-clock duration spanning replacement provisioning, image pull, weight loading, and readiness gate passage.
  • llm_spot_wasted_decode_tokens_total: Total volume of generated decode tokens discarded due to interrupted runs.

Chaos Engineering Drill Matrix

Prior to production launch, use AWS Fault Injection Service (FIS) or local ACPI event injection tools to validate system resilience:

  • Scenario 1: Trigger preemption while nodes run at 80% VRAM utilization with high-concurrency, short-sequence traffic; measure the drain completion success rate.
  • Scenario 2: Inject an interruption during a 4K+ token prefill phase; verify prompt abort triggers and rapid failover execution.
  • Scenario 3: Simulate regional Spot exhaustion; verify that the autoscaler falls back to On-Demand capacity within target SLO limits.
  • Scenario 4: Terminate a connection mid-stream; confirm that downstream clients consume error codes correctly without triggering infinite retry storms.

FAQ

What is the very first action to take when a Spot GPU receives an interruption notice?
Immediately shut down ingress admission by setting the readiness probe to false and detaching the node at the gateway level. Mark the instance as Draining, and trigger warm replacement provisioning in parallel with draining. Never wait for inflight inference requests to finish before requesting replacement capacity.
Can streaming responses (SSE/WebSocket) interrupted by Spot preemption be retried directly?
Do not retry blindly. In streaming contexts, the client has typically already received partial tokens; recomputing from scratch without state tracking causes duplicate content or semantic drift. Pure read-only requests should be managed using a gateway request ledger with prefix continuation. Workflows involving writes, billing, or tool calling must enforce strict idempotency keys.
Does a PodDisruptionBudget (PDB) prevent cloud providers from abruptly reclaiming Spot instances?
No. PDBs only govern voluntary evictions initiated by the Kubernetes API. They cannot prevent involuntary, physical infrastructure reclamation by cloud providers. A PDB controls the maximum unavailable replicas during planned cluster operations, but offers no guarantee against Spot preemptions.
Can we achieve seamless failover in production by migrating the KV Cache across nodes?
Not recommended as a general disaster-recovery prerequisite. Synchronizing VRAM state across nodes via RDMA introduces massive operational complexity and demands strictly homogeneous hardware topologies, yielding low ROI within a brief 30-to-120-second eviction window. Production architectures should assume node loss equals total KV cache loss, relying instead on drain budgets, early aborts, and tiered capacity pools.