AIAny
AI Model2026
Icon for item

Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw

Provides an EXL3 3.0 bits-per-weight quantization of a weight-edited GLM-5.3 UNCENSORED FP8 model for self-hosted text generation and agent workflows. Key characteristics: 753B MoE architecture, 273 GiB on disk, converted with ExLlamaV3; tool-call parsing requires preserving string arguments.

Introduction

This release makes a weight-edited, alignment-stripped variant of GLM-5.3 feasible to self-host by compressing it into an EXL3 3.0 bpw format — but the savings come with concrete tradeoffs users must accept and mitigate.

Key Capabilities
  • Large MoE at reduced storage: the quant stores a 753B-parameter GlmMoeDsaForCausalLM as a 273 GiB EXL3 pack (avg 3.04 bpw), enabling multi-GPU inference with ExLlamaV3 and TabbyAPI. So what: you can run GLM-5.3–class behavior without full-precision costs, but you still need a substantial GPU cluster.
  • Fidelity and measurable deltas: small degradation versus the FP8 source (KL divergence ~0.089, perplexity 3.440 vs 3.302). So what: most high-confidence token predictions are nearly unchanged, but some downstream agent tasks show modest differences.
  • Agent and tool-call behaviour: tested on 8×A100(40GB) with gpu_split_auto, MTP drafting and 98K shared cache; the model emits tool arguments as raw-tagged text, which can be JSON-decoded into wrong types by naive parsers. So what: production agent pipelines must validate or preserve string-typed parameters to avoid failed tool calls.
Who it's for and tradeoffs

Great fit if you need an uncensored/weight-edited GLM-5.3 variant for local or rented multi-GPU hosting and are prepared to handle agent parsing edge cases, large disk footprint (273 GiB) and license/ethics responsibilities. Look elsewhere if you need turnkey, safety-aligned hosted inference, single-GPU convenience, or strict guarantees about unchanged vendor behavior — this build applies a permanent weight edit (no fine-tune) and is community-converted, not an upstream vendor release.

Where it fits

Positioned between full-precision vendor FP8 deployments and smaller distilled models: it prioritizes fidelity to a specific weight-edited variant while lowering storage via EXL3 quantization. Use it for experimentation, self-hosted agent research, or environments where you control tool parsing and infrastructure; avoid it for public-facing, safety-critical services without additional safeguards.

Information

  • Websitehuggingface.co
  • OrganizationsInfatoshi, dealignai, zai-org
  • AuthorsInfatoshi
  • Published date2026/10/01

Categories

More Items

Hugging Face
AI Model2026

Processes English and German text with long-context reasoning and structured tool-calling. Uses a 78B mixture-of-experts architecture that activates ~3.46B parameters per token, offers native 262k-token context (validated to 1M), and is released as Apache-2.0 weights — suited for RAG, document processing and human-in-the-loop decision support.

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.