A single-file, reproducible recipe for the exact config I run in production: full GLM-5.2
Int4-Int8Mix → 5% expert-prune → NVFP4 4-bit KV cache → MTP=3 speculative decode → 100K context,
served as glm-5.2-full on a 4× DGX Spark / GB10 cluster (TP=4). All numbers are greedy (temp=0).
Headline: ~9.5 t/s (full, eager) → ~21.4 t/s (5%-prune + MTP=3) at 100K context, ~360 GB/node, coding 5/5, Ukrainian long-form clean (10% prune already breaks it — 5% is the safe ceiling).