AIAny
AI Model2026
Icon for item

orcarouter/Qwen3.8-27B-Uncensored-MLX

Provides a quantized MLX conversion of Qwen3.8-27B for Apple Silicon (2/4/6/8-bit) with the model's refusal-direction ablated, preserving multimodal vision+text capability; intended for red‑teaming, interpretability and safety research, not unmoderated production use.

Introduction

Why this matters

This Hugging Face repository supplies a post-quantized MLX build of Qwen3.8-27B where the model's refusal direction has been ablated — i.e., the built-in safety/refusal behavior is substantially removed. That makes the artifact valuable for controlled red‑teaming, interpretability, and guardrail research because probes that the original model would refuse now return substantive outputs, while keeping the original model's multimodal capabilities intact.

What Sets It Apart
  • Abliteration of refusal behavior: a directional ablation was applied to orthogonalize out refusal signals from the LM residual stream, producing a model that returns substantive answers on red‑team probes where the base model would refuse.
  • MLX quantizations for Apple Silicon: four precision builds are provided (2 / 4 / 6 / 8-bit, MLX affine, group size 64). The repo root mirrors the 4-bit build for easy loading; 4/6/8-bit are recommended for usable quality, 2-bit is archival and severely degraded.
  • Multimodal fidelity preserved: the vision tower and normalization layers are left in BF16, so image understanding (shapes, colors, layout, text-in-image) is retained while language linear weights are quantized.
  • Long context and native VLM architecture: preserves Qwen3.8 architecture traits (hybrid Gated DeltaNet + attention, 64 text layers) and a 262,144-token context window.
Who it's for and tradeoffs

Great fit if you want a locally runnable, uncensored derivative of Qwen3.8-27B for: red‑teaming, guardrail/robustness evaluation, refusal‑mechanism studies, and interpretability experiments. The upload includes practical quantization choices (4-bit default ≈15 GB; 6/8-bit for higher fidelity) and tradeoffs are explicit: the model intentionally lacks meaningful built-in safety guardrails, so outputs may be harmful, illegal, or offensive. Users must add their own moderation and abuse-prevention layers before any deployment. The base model credit and Apache‑2.0 license remain; the uploader disclaims liability for misuse.

Technical notes (concise)

  • Available precisions: 8-bit (~27.5 GB), 6-bit (~22 GB), 4-bit (~15 GB, repo root), 2-bit (~8.7 GB, archival). Recommended: 4/6/8-bit; avoid 2-bit for real tasks.
  • Kept in BF16: vision tower, all norms, and certain conv/attention tensors; Quantized: language linear layers incl. embed_tokens and lm_head.
  • Intended usage: controlled research environments only; not for production-facing or unmoderated deployments.

Information

  • Websitehuggingface.co
  • Organizationsorcarouter (uploader), Qwen / Alibaba Cloud (base model)
  • Authorsorcarouter
  • Published date2026/08/17

Categories

More Items

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.