This release makes a weight-edited, alignment-stripped variant of GLM-5.3 feasible to self-host by compressing it into an EXL3 3.0 bpw format — but the savings come with concrete tradeoffs users must accept and mitigate.
Key Capabilities
- Large MoE at reduced storage: the quant stores a 753B-parameter GlmMoeDsaForCausalLM as a 273 GiB EXL3 pack (avg 3.04 bpw), enabling multi-GPU inference with ExLlamaV3 and TabbyAPI. So what: you can run GLM-5.3–class behavior without full-precision costs, but you still need a substantial GPU cluster.
- Fidelity and measurable deltas: small degradation versus the FP8 source (KL divergence ~0.089, perplexity 3.440 vs 3.302). So what: most high-confidence token predictions are nearly unchanged, but some downstream agent tasks show modest differences.
- Agent and tool-call behaviour: tested on 8×A100(40GB) with gpu_split_auto, MTP drafting and 98K shared cache; the model emits tool arguments as raw-tagged text, which can be JSON-decoded into wrong types by naive parsers. So what: production agent pipelines must validate or preserve string-typed parameters to avoid failed tool calls.
Who it's for and tradeoffs
Great fit if you need an uncensored/weight-edited GLM-5.3 variant for local or rented multi-GPU hosting and are prepared to handle agent parsing edge cases, large disk footprint (273 GiB) and license/ethics responsibilities. Look elsewhere if you need turnkey, safety-aligned hosted inference, single-GPU convenience, or strict guarantees about unchanged vendor behavior — this build applies a permanent weight edit (no fine-tune) and is community-converted, not an upstream vendor release.
Where it fits
Positioned between full-precision vendor FP8 deployments and smaller distilled models: it prioritizes fidelity to a specific weight-edited variant while lowering storage via EXL3 quantization. Use it for experimentation, self-hosted agent research, or environments where you control tool parsing and infrastructure; avoid it for public-facing, safety-critical services without additional safeguards.