Why this matters
By trimming routed experts and storing expert weights in NVFP4, this build makes a GLM-5.3-Flash–class multimodal model practical to run on desktop Blackwell hardware. The core insight is memory-first optimization: reduce resident expert footprint enough to leave meaningful KV cache on each DGX Spark node while keeping the original model’s active-parameter profile and accuracy class.
Key capabilities
- Memory-optimized GLM-5.3-Flash derivative: keeps 224 of 288 routed experts per layer and stores experts in modelopt NVFP4 (16-element groups, e4m3 group scale, fp32 tensor scale), yielding ~141 GiB of weights (151.5 GB on disk).
- Multimodal and compatibility: retains the vision tower and full 154,880-token vocabulary; validated for image+text use cases and tool calling in vLLM.
- Production-oriented serving: validated on a single NVIDIA B200 (GB200) and provisioned for 2× DGX Spark (ConnectX‑7) deployments; includes an optional MTP speculative-decoding draft that can ≈1.85× single-stream decode speed in low-concurrency interactive serving.
- Performance profile: still activates ~18B parameters per token (top-8 of 224 experts + shared expert + attention); measured benchmarks on a B200 show HumanEval ~98.2% (sampling), GPQA‑Diamond 90.9% at max effort, AIME 2025 pass@1 88.3%.
- Deployment notes: recommended vLLM ≥ 0.30.0 (or vllm/vllm-openai:glm53-flash image), transformers ≥ 5.16.1, and VLLM_USE_DEEP_GEMM=0. Memory budgeting guidance and a 2× DGX Spark runbook are provided in the repo.
Who it's for, and tradeoffs
Great fit if you need a near–full-quality GLM-5.3-Flash experience on Blackwell desktop class hardware and can provision 2× DGX Sparks or a single ≥180 GB B200. It’s tailored to interactive, low-to-moderate concurrency serving where KV cache size and memory bandwidth matter (MTP is recommended for single-user interactive latency). Look elsewhere if you require a single 128 GB Spark deployment (weights exceed 128 GB) or absolute parity with the unpruned FP8 checkpoint in every micro-benchmark; this is an unofficial, community-built derivative and may have small accuracy/behavior gaps compared with vendor-published checkpoints.