LLM Training NaN/Inf Production Guide: Non-finite Gates, GradScaler Evidence, and Layer Bisect to Isolate the First Bad Step
A practical guide to handling NaN/Inf in large model training. Covers non-finite gates, GradScaler skip evidence, distributed consistent skipping, layer bisect localization, and bad batch isolation for replayable, stoppable, and recoverable numerical anomaly management.