AIAny
AI Model2026
Icon for item

OrcaSAQ-2-Cyber-27B-Uncensored-GGUF

Provides a high-fidelity mixed-precision (≈3-bit) GGUF quant of Qwen3.8-27B tailored for long-horizon agent workloads and cyber-focused red-teaming. Preserves reasoning, thinking mode, MTP speculative decoding and vision (via a separate mmproj); released as an uncensored/abliterated research build under Apache-2.0.

Introduction

Making a 27B reasoning-capable model practical for long-horizon agentic workflows and specialist cyber evaluations is the core trade-off this release targets: it compresses a full Qwen3.8-27B checkpoint into a GPU-friendly GGUF quant while intentionally exposing an "uncensored" behavior profile for interpretability and red-teaming.

Key Capabilities
  • Extreme compression with high fidelity: OrcaSAQ2-style quantization reduces the original BF16 checkpoint (54 GB) to roughly 12.3 GB while reporting only +0.02% perplexity, ~93.2% token-level top-1 agreement, and mean KLD ≈ 0.031 on calibration tests — preserving long-context (262k) and reasoning behavior.
  • Architecture and feature preservation: the hybrid Gated DeltaNet + full-attention architecture, the MTP speculative-decoding head, thinking mode and tool-calling capabilities are retained; vision support is provided via a separate mmproj file and the build targets llama.cpp / GGUF runtimes.
  • Research and red-team orientation: the build is explicitly "abliterated" (refusal directions suppressed) to study refusal mechanisms, robustness, and offensive-security capabilities; users are expected to add their own safety and moderation layers before any deployment.
Who it's for and tradeoffs

Great fit if you are doing offline research, interpretability or red-teaming on large multimodal LLMs and need a compact, long-context-capable 27B that still supports thinking/tooling primitives. Look elsewhere if you need a production-safe model out-of-the-box: the uncensored nature removes many safety refusals and increases misuse risk. Additional tradeoffs include modest fidelity degradation at extreme low-bit quants and the requirement of a recent llama.cpp build (qwen35 + nextn/MTP support) or compatible runtimes to fully restore architecture and speculative decoding behavior.

In short: a practical, high-fidelity quantized path to run a reasoning-capable Qwen3.8-27B variant for long-horizon agents and cyber-focused research, but not a turnkey solution for public-facing production without substantial safety controls.

More Items

Hugging Face
AI Model2026

Capability-targeted compression of Qwen3.8-Flash-Next: half the experts are removed and remaining weights quantized to 3.5 bpw, producing a 58.4 GB GGUF (29.6 GB resident) that preserves coding and multimodal ability while trading off other domains.

Hugging Face
AI Model2026

Provides compact mixed-precision GGUF quantizations of UkisAI's Swift 1.5 (derived from Qwen3.8-27B), using ISTA GSQ-RCO per-tensor allocations with Swift-specific refinement. Offers multiple 8–12 GB tiers, optional MTP heads, and KLD evaluation against the Swift BF16 baseline.

Hugging Face
AI Model2026

Provides official pretrained VisionHOPE visual-backbone checkpoints for ImageNet classification, COCO object detection & instance segmentation, and ADE20K semantic segmentation. Includes hierarchical Tiny/Small/Base models with PyTorch-compatible downloadable weights.