Four Little Boxes = One Frontier Model on DGX Sparks

There’s a particular satisfaction in watching a model that has no business running on the hardware in front of you generate its first coherent token. We got GLM-5.2 — the full, unpruned model, all 256 experts — serving at 327,000 tokens of context at about 25 tokens per second, across four NVIDIA DGX Spark boxes […]