AIAny
AI Model2026
Icon for item

Qwen3.8-27B-Uncensored-FP8

An FP8-quantized, uncensored mirror of Qwen3.8-27B for image-text-to-text tasks — preserves native multimodal vision and very long context while targeting transformers/vLLM deployments; intended for offline testing and red-teaming and may bypass built-in safety filters.

Introduction

Why this matters Qwen3.8-27B-Uncensored-FP8 makes a deployment-friendly, quantized copy of a 27B multimodal causal model available for offline testing and red-teaming. By combining FP8 compression with the Qwen3.8 architecture’s native vision and very long context support, it lowers resource barriers for experimenting with agentic and long-horizon multimodal workflows — at the cost of removing the original model’s content refusals.

Key Capabilities
  • FP8 quantization for linear layers: reduces storage and memory footprint compared with full-precision checkpoints, making single-GPU or constrained-cluster deployment more accessible while keeping output heads and some vision/attention components at full precision for stability.
  • Native multimodal support: retains the Qwen3.8 family’s image/video understanding and image->text pipeline suitability, useful for document, diagram, and visual reasoning tasks.
  • Very long context and agentic features: retains the native 262,144-token context window (extensible via YaRN techniques) and the Qwen3.8 “thinking” controls (e.g., reasoning depth tuning/preserved thinking) that help with multi-step planning and chain-of-thought style traces.
  • Compatibility: packaged for Hugging Face Transformers/safetensors and commonly used inference stacks such as vLLM and other compressed-tensor toolchains, enabling easy integration into existing evaluation and deployment pipelines.
Who it's for & tradeoffs

Great fit if you need a locally deployable, multimodal 27B model for offline evaluation, red-teaming, or resource-constrained inference workloads and you accept reduced precision in exchange for smaller checkpoints and memory use. Look elsewhere if you require out-of-the-box safety/refusal behavior, strict compliance with content-moderation policies, or absolute bit-for-bit parity with the original full-precision model — FP8 quantization and the “uncensored” nature change numerical behavior and content controls. Also evaluate numerical stability on your target tasks: some attention/visual layers and the LM head are commonly kept at higher precision to preserve fidelity, but behavior can still differ from full-precision checkpoints.

More Items

Hugging Face
AI Model2026

Compresses Qwen3.8-27B into a 12.3 GB sensitivity-aware mixed-precision quantized checkpoint for long-horizon agent workloads; preserves BF16 fidelity (+0.02% PPL, 93.2% token Top‑1 agreement), supports 262K context and vLLM serving, text-only and Apache‑2.0 licensed.

Hugging Face
AI Model2026

Turns a context, a question, and 2–20 candidate answers into a single chosen option for classification, routing, ordered scores, and Boolean decisions. Built on mmBERT-small with a 144.3M-parameter decision head, runs on CPU with up to 8,192 combined tokens; FP32 weights occupy 550.5 MiB and are Apache‑2.0 licensed.

Hugging Face
AI Model2026

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.