AIAny
AI Model2026
Icon for item

huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF

An uncensored, "abliterated" fork of Qwen3.8-27B that removes refusal behavior by modifying targeted weights and provides multiple GGUF/BF16 quantized variants for local research and deployment, while carrying significantly reduced safety filtering.

Introduction

Altering a model's refusal behavior changes how it handles sensitive prompts and can dramatically widen possible outputs — intentionally or not. This release is a proof-of-concept "abliteration" of Qwen3.8-27B that modifies specific tensors to remove safety refusals and then offers several quantized builds for local use and experimentation.

What Sets It Apart
  • Abliteration approach: specific weight tensors (token_embd, output, ffn_down, ssm_out, attn_output) are targeted and converted/modified to remove refusal behavior rather than retraining or using activation-space edits. The first 15 transformer layers, MTP, and visual components are retained unmodified.
  • Multi-format distribution: provides BF16 base plus multiple GGUF quantized builds (Q2_K_L through Q8_0_L variants, with some weights converted to BF16 and filenames suffixed _L) to balance size, speed, and quality for different local runtimes.
  • Local-runtime compatibility: prepared for common local inference tools (llama.cpp, ollama) and includes guidance for llama-quantize conversion commands so users can run appropriate GGUF variants on CPU/GPU environments.
  • Research-first posture: explicitly described as experimental and uncensored; maintainers warn about reduced safety filtering and recommend monitoring outputs and restricted use cases.
Who it's for and tradeoffs

Great fit if you want a research-oriented, local copy of Qwen3.8 modified to bypass refusal behavior for experimentation, forensic analysis of safety mechanisms, or studying the effect of targeted weight changes across quantization formats. It saves time for users who need quantized GGUF builds ready for llama.cpp/ollama.

Look elsewhere if you need production-ready, safety-hardened models or are building public-facing services; the project reduces safety filtering and carries ethical, legal, and content-moderation risks. Expect unpredictable behaviors and the need for manual output review; not suitable for minors or uncontrolled public deployment.

Information

Categories

More Items

Hugging Face
AI Model2026

Provides an EXL3 3.0 bits-per-weight quantization of a weight-edited GLM-5.3 UNCENSORED FP8 model for self-hosted text generation and agent workflows. Key characteristics: 753B MoE architecture, 273 GiB on disk, converted with ExLlamaV3; tool-call parsing requires preserving string arguments.

Hugging Face
AI Model2026

Processes English and German text with long-context reasoning and structured tool-calling. Uses a 78B mixture-of-experts architecture that activates ~3.46B parameters per token, offers native 262k-token context (validated to 1M), and is released as Apache-2.0 weights — suited for RAG, document processing and human-in-the-loop decision support.

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.