AIAny
AI Model2026
Icon for item

GLM-5.2 (Unsloth GGUF)

Provides GGUF-quantized GLM-5.2 builds for local text-generation with a solid 1M-token context, dynamic 1-/2-bit quant options, and Unsloth runtime integrations — targeted at long-horizon coding, reasoning and agent workflows. MIT licensed.

Introduction

Long-context LLMs are difficult to run locally at scale; this GGUF distribution packages GLM-5.2 so you can run the model with Unsloth dynamic quantization and common local inference tooling.

Key Capabilities
  • 1M-token context: Enables stable long-horizon tasks (editing, long-form reasoning, multi-file codebases) without fragmenting context.
  • Quantized GGUF builds: Dynamic 1-bit and 2-bit variants (and higher-bit options) let you trade model footprint versus fidelity; 1-bit ≈ 223 GB total memory, 2-bit ≈ 239 GB on disk in common distributions.
  • Runtime & integration: Designed for llama.cpp, Unsloth Studio, vLLM and transformers ecosystems; includes presets for reasoning effort (non-thinking, high, max) and speculative decoding improvements.
  • Architecture notes: Builds on GLM-5.2 innovations (IndexShare sparse indexing and MTP improvements) to reduce per-token FLOPs at very long contexts and increase speculative decoding acceptance.
Who it's for and trade-offs

Great fit if you want to run a large long-context LLM locally or on-premise (researchers, teams testing agentic chains, developers evaluating long-form code generation) and can provide large unified memory or a mix of VRAM+RAM. Look elsewhere if you need a tiny footprint (edge devices) or cannot meet the hundreds of GBs of total memory required for useful quantized variants. The package prioritizes reproducible, local inference and measurable trade-offs between quant levels (file size vs. accuracy).

Information

  • Websitehuggingface.co
  • Organizationsunsloth, Z.ai (zai-org/GLM-5)
  • Published date2026/06/17

Categories

More Items

Hugging Face
AI Model2026

Compresses Qwen3.8-27B into a 12.3 GB sensitivity-aware mixed-precision quantized checkpoint for long-horizon agent workloads; preserves BF16 fidelity (+0.02% PPL, 93.2% token Top‑1 agreement), supports 262K context and vLLM serving, text-only and Apache‑2.0 licensed.

Hugging Face
AI Model2026

Turns a context, a question, and 2–20 candidate answers into a single chosen option for classification, routing, ordered scores, and Boolean decisions. Built on mmBERT-small with a 144.3M-parameter decision head, runs on CPU with up to 8,192 combined tokens; FP32 weights occupy 550.5 MiB and are Apache‑2.0 licensed.

Hugging Face
AI Model2026

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.