LLM GPU Memory Fragmentation in Production: Taming OOMs with Memory Snapshots, Allocator Tuning, and Admission Control
LLM inference services often hit OOM due to fragmentation, not total capacity exhaustion. This guide covers Memory Snapshots, dynamic shapes, allocator tuning, and vLLM memory budgeting for observable, verifiable, and rollback-safe production management.