LLM Knowledge Distillation in Production: Controlling Student Model Degradation with Teacher Logprob Contracts, On-Policy Replay, and Top-K Tail Buckets
A practical engineering guide to productionizing LLM knowledge distillation, covering Teacher Logprob contracts, temperature and divergence selection, On-policy replay, Top-K tail probability buckets, Tokenizer alignment, and student capability gating to prevent silent degradation and deployment failures.