Tag: Tail Latency

2 articles

  • LLM capacity testing shouldn't rely on fixed concurrency and average latency alone. This article combines vLLM, AIPerf, and MLPerf to explain Timed Traces, burst traffic, tail-latency gates, and Safe QPS — a production-grade method for finding the true capacity inflection point before launch.

  • Long contexts and agentic requests shift LLM inference bottlenecks from single-card throughput to tail-latency governance. This post dives into Prefill-Decode disaggregation, KV cache transfer, resource planning, and a pre-launch checklist to move LLM serving from 'working' to stable and scalable.