Why Are LLMs More Vulnerable to Spot Preemption Than Standard Services?
Spot and preemptible GPU instances typically deliver 60% to 80% cost savings compared to On-Demand instances, making them the primary lever for reducing LLM inference infrastructure bills. However, directly applying generic stateless web app Spot practices to LLM serving frameworks (such as vLLM, TGI, or TensorRT-LLM) often triggers catastrophic availability drops:
- High Execution Time: The prefill phase for large prompts and the decoding loop for hundreds to thousands of tokens can run from tens of seconds to several minutes. When an interruption signal arrives, nodes almost always hold deep backlogs of long-running, half-finished requests.
- In-Memory KV Cache Residency: Self-attention KV caches reside entirely within VRAM. A hard termination implies that all intermediate computational state is lost; cross-node zero-downtime checkpoints are practically nonexistent.
- Irreversible Streaming Tokens: When using Server-Sent Events (SSE) or WebSockets, earlier tokens have already been delivered to the client. If an interrupted request is restarted from scratch on a new node, it readily leads to duplicate responses, branching outputs, or semantic drift. If tool use or function calling is involved, it can trigger duplicate external side effects.
Spot GPU management should not be built on the hope that preemptions won’t happen. Instead, you need a deterministic control pipeline: Signal Ingestion → Admission Close → Drain Budget → Warm Replacement → Idempotent Retry.
Principle 1: Converting the Notice Window into a Drain Budget
Cloud providers offer varying early termination notices, all provided strictly on a best-effort basis:
| Cloud / Platform | Signal Type | Typical Window | Behavioral Profile |
|---|---|---|---|
| AWS EC2 Spot | Interruption Notice / Rebalance Recommendation | ~120s (Notice) / Earlier (Recommendation) | Rebalance signals carry no SLA; instances are hard-terminated at 120s |
| Google Cloud (GCP) | Metadata Preemption Notice | ~30s (Default) / Configurable (Preview) | Sends an ACPI shutdown signal with an extremely tight default window |
| Kubernetes (Karpenter) | Node Interruption Event | Derived from underlying cloud events | Applies taints automatically and initiates a graceful node drain |
In production, you cannot assume you have the entire warning duration. You must calculate the actual Drain Budget by subtracting end-to-end control-plane latency:
$$\text{DrainBudget} = \text{NoticeWindow} - \text{DetectionLatency} - \text{IngressRemovalLatency} - \text{SafetyMargin}$$
- DetectionLatency: Delay from DaemonSet metadata polling or cloud EventBridge notifications (keep this under 1–3 seconds).
- IngressRemovalLatency: Time required for EndpointSlice updates, Kube-Proxy syncs, and load balancer/ingress deregistration (typically 2–8 seconds).
- SafetyMargin: Buffer for network jitter, context flushing, and SIGKILL boundaries (recommend reserving 10–15 seconds).
On AWS with a 120-second window, your actionable Drain Budget is usually 85–95 seconds. Under GCP’s default 30-second window, your realistic budget drops below 15 seconds.
Principle 2: Parallelize Warm Replacement and Draining — Stop Waiting Sequentially
A common anti-pattern in cluster operations is running lifecycle steps sequentially: receive termination signal → drain old Pod → wait for Pod exit → trigger autoscaler scale-out → pull container image and weights → register new Pod. With LLMs, cold starts frequently take several minutes; running these phases serially causes an immediate collapse of cluster serving capacity.
The production-ready architecture splits execution into two parallel paths immediately upon signal confirmation:
┌── 1. Close admission + race inflight requests within Drain Budget (Old Node)
[Interruption Signal / Taint] ────┤
└── 2. Request replacement instance + warm replacement in parallel (New Node)
By leveraging Karpenter interruption handling or custom autoscaler controllers, you can schedule replacement capacity the instant an evicted node receives a cloud.google.com/gke-preemptible=true or karpenter.sh/disruption=interrupted taint. This maximizes the overlap between bringing the new node up and draining the old node.
Principle 3: Classify Inflight Requests to Cut Losses Early
Once an instance enters the Draining state, it should never passively wait for all requests to finish. Instead, use an estimation algorithm based on remaining tokens and current throughput to decide whether an inflight request can finish within the allocated budget:
def evaluate_inflight_request(req, drain_budget, recent_tpot, safety_factor=0.85):
"""
Decide whether to finish processing the request locally based on estimated time.
"""
estimated_remaining_tokens = req.max_tokens - req.generated_tokens
estimated_remaining_time = estimated_remaining_tokens * recent_tpot
# Continue only if execution fits safely within the discounted Drain Budget
if estimated_remaining_time <= (drain_budget * safety_factor):
return "CONTINUE_DRAIN"
else:
return "ABORT_AND_FAILOVER"
Route inflight requests based on their operational profiles:
| Request Profile | Evaluation Criteria | Mitigation Action |
|---|---|---|
| Short outputs / Late decoding | Est. completion time $\le$ Remaining Drain Budget | Continue decoding and complete before eviction |
| Long outputs / Early prefill | Est. completion time $>$ Remaining Drain Budget | Abort immediately to prevent wasted compute |
| Read-only idempotent inference | Client supports retries or gateway manages state | Intercept and reroute to a healthy warm node |
| Composite / Side-effecting | Writes to DB, processes payments, or runs Tool Calls | Record state in gateway ledger; block auto-replay and return explicit status |
Principle 4: Streaming Retry Contracts and Gateway Ledgers
Non-streaming JSON requests can be retried simply by verifying idempotency keys. Streaming responses, however, present a major UX challenge: clients have often already received and displayed initial tokens. Starting over from the raw prompt on a new node causes UI flickering, duplicated sentences, and broken experiences.
We recommend deploying a context-aware request ledger inside the API Gateway (e.g., via Envoy, Kong, or a custom ingress proxy):
{
"request_id": "req_llm_9837421",
"attempt": 2,
"served_prefix_tokens": 214,
"model_revision": "qwen2.5-72b-instruct@v1",
"prompt_fingerprint": "sha256_e3b0c44...",
"retry_safe": true,
"interruption_reason": "spot_eviction"
}
Retry Protocol Strategies:
- Clean Break Protocol (Client-Managed): The gateway emits a structured SSE error event:
event: error, data: {"code": "SPOT_PREEMPTION", "recoverable": true}. The client-side application captures this event, alerts the user, and automatically sends an append-mode request with the previously rendered text acting as the conversational prefix. - Server-side Prefix Continuation: If the serving engine supports prefix caching and the prompt uses deterministic decoding parameters (
temperature=0, fixed seed), the gateway can inject the previously delivered tokens as a prefix into the new replica, streaming back only the delta. This requires strict model version and sampling consistency.
Architecture in Practice: End-to-End Spot Preemption State Machine
+-------------+ Interruption Warning +--------------+
| READY | ------------------------> | DRAINING |
+-------------+ +--------------+
^ | |
| | | (Parallel Actions)
| v v
+-------------+ Model Gate Passed [Stop Ingress] [Provision Replacement]
| RECOVERED | <------------------ | |
+-------------+ [Run or Abort] [Warm Weights & GPU]
| |
+----+-----+
|
v
+--------------+
| TERMINATING |
+--------------+
Critical State Transitions:
- Signal Interception (Node-Problem-Detector / Spot Handler): Continuously poll instance metadata. The moment a preemption event is discovered, apply a
NoScheduletaint to the Kubernetes node and flip the serving Pod’s readiness status tofalse. - Application Drain Loop:
async def handle_spot_eviction(deadline_ts):
serving_engine.stop_accepting_new_requests()
while get_current_time() < (deadline_ts - SAFETY_MARGIN_SECONDS):
if serving_engine.get_active_request_count() == 0:
break
await asyncio.sleep(0.5)
# Abort requests that cannot finish within the budget
await serving_engine.abort_all_inflight(
reason="drain_budget_exhausted_spot_cleanup"
)
- Strict Replacement Readiness Gates: Before routing traffic to the replacement instance, never rely solely on standard
PodScheduledorContainersReadyconditions. Enforce an end-to-end model verification probe:- Validate GPU driver status and memory allocation.
- Confirm weights are deserialized and bound to device VRAM.
- Run a dummy prompt execution to warm up CUDA kernels and execution graphs.
Compute Topology: Hybrid Pools and Failure Domain Isolation
Never rely exclusively on a single GPU instance type within a single Availability Zone for your Spot inference workloads. A regional supply squeeze can wipe out your entire pool simultaneously, leading to an availability cascade.
┌── On-Demand Baseline Capacity (Sustains 20%–40% Core Load)
│
Inference Cluster ──┼── Spot Pool A (Primary AZ, e.g., A100-80G / AZ-a)
├── Spot Pool B (Alternative Instance/AZ, e.g., L40S / AZ-b)
└── Fallback Degradation Pipeline (Load shedding / Quantized model fallback)
- Baseline Floor: Keep an On-Demand compute baseline dedicated to your minimum viable SLO. If all Spot instances are revoked simultaneously, your service degrades gracefully rather than suffering an outright outage.
- Instance Diversification: Broaden Karpenter node pool requirements. Allow the scheduler to pick across equivalent GPU families (e.g., fall back across A100, H100, L40S, or A10G depending on model compatibility).
Golden Metrics and Chaos Engineering Checklist
Core Prometheus Metrics
llm_spot_notice_to_draining_seconds: Latency between cloud provider interruption emission and the application closing ingress admission.llm_spot_inflight_aborted_ratio: Percentage of requests forcefully terminated due to exhausted drain budgets.llm_spot_replacement_ready_latency_seconds: Wall-clock duration spanning replacement provisioning, image pull, weight loading, and readiness gate passage.llm_spot_wasted_decode_tokens_total: Total volume of generated decode tokens discarded due to interrupted runs.
Chaos Engineering Drill Matrix
Prior to production launch, use AWS Fault Injection Service (FIS) or local ACPI event injection tools to validate system resilience:
- Scenario 1: Trigger preemption while nodes run at 80% VRAM utilization with high-concurrency, short-sequence traffic; measure the drain completion success rate.
- Scenario 2: Inject an interruption during a 4K+ token prefill phase; verify prompt abort triggers and rapid failover execution.
- Scenario 3: Simulate regional Spot exhaustion; verify that the autoscaler falls back to On-Demand capacity within target SLO limits.
- Scenario 4: Terminate a connection mid-stream; confirm that downstream clients consume error codes correctly without triggering infinite retry storms.