AIAny
AI Model2026
Icon for item

GLM5.3-Flash-E224-DGX-Spark

Runs a pruned, NVFP4-quantized GLM-5.3-Flash variant tuned for Blackwell GPUs: 224 routed experts per layer and ~141 GiB of weights. Retains the multimodal vision tower, activates 18B params/token, supports vLLM and optional MTP speculative decoding; fits 2× DGX Spark or a ≥180 GB B200.

Introduction

Why this matters

By trimming routed experts and storing expert weights in NVFP4, this build makes a GLM-5.3-Flash–class multimodal model practical to run on desktop Blackwell hardware. The core insight is memory-first optimization: reduce resident expert footprint enough to leave meaningful KV cache on each DGX Spark node while keeping the original model’s active-parameter profile and accuracy class.

Key capabilities
  • Memory-optimized GLM-5.3-Flash derivative: keeps 224 of 288 routed experts per layer and stores experts in modelopt NVFP4 (16-element groups, e4m3 group scale, fp32 tensor scale), yielding ~141 GiB of weights (151.5 GB on disk).
  • Multimodal and compatibility: retains the vision tower and full 154,880-token vocabulary; validated for image+text use cases and tool calling in vLLM.
  • Production-oriented serving: validated on a single NVIDIA B200 (GB200) and provisioned for 2× DGX Spark (ConnectX‑7) deployments; includes an optional MTP speculative-decoding draft that can ≈1.85× single-stream decode speed in low-concurrency interactive serving.
  • Performance profile: still activates ~18B parameters per token (top-8 of 224 experts + shared expert + attention); measured benchmarks on a B200 show HumanEval ~98.2% (sampling), GPQA‑Diamond 90.9% at max effort, AIME 2025 pass@1 88.3%.
  • Deployment notes: recommended vLLM ≥ 0.30.0 (or vllm/vllm-openai:glm53-flash image), transformers ≥ 5.16.1, and VLLM_USE_DEEP_GEMM=0. Memory budgeting guidance and a 2× DGX Spark runbook are provided in the repo.
Who it's for, and tradeoffs

Great fit if you need a near–full-quality GLM-5.3-Flash experience on Blackwell desktop class hardware and can provision 2× DGX Sparks or a single ≥180 GB B200. It’s tailored to interactive, low-to-moderate concurrency serving where KV cache size and memory bandwidth matter (MTP is recommended for single-user interactive latency). Look elsewhere if you require a single 128 GB Spark deployment (weights exceed 128 GB) or absolute parity with the unpruned FP8 checkpoint in every micro-benchmark; this is an unofficial, community-built derivative and may have small accuracy/behavior gaps compared with vendor-published checkpoints.

More Items

Hugging Face
AI Model2026

Performs a byte-level transplant of 144 tensors in an already-quantized GSQ-RCO Qwen3.8-Flash-Next to ablate the model's refusal direction while preserving GSQ-learned scales and the upstream per-tensor type assignment; multimodal, 262K context. Intended for local inference, red-teaming and quantization research; no retraining or built-in safety.

Hugging Face
AI Model2026

Maps multimodal inputs (text + images) to structured decisions (yes/no, choice, or scored rubric) in a single forward pass and returns calibrated probabilities. 3.1B parameters, long context (32,768 tokens), optimized for low-latency edge inference; not a text-generation/chat model.

Hugging Face
AI Model2026

Returns calibrated probabilities for yes/no and multi-choice decisions using a two-stage System 1 (fast classifier) and System 2 (Gemma‑4 reasoning) pipeline; this NVFP4 release quantizes the 3,840 routed MoE experts to 4-bit so the model fits ≈17–18 GB of GPU memory while other weights remain bf16.