AIAny
AI Model2026
Icon for item

Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF

Provides quantized GGUF variants of Qwen3.8-27B with an 'Aggressive' uncensoring profile and an optional HauhauCS FastMTP sidecar to accelerate MTP speculative decoding; includes a BF16 vision projector and K_P quant levels for VRAM/quality trade-offs.

Introduction

This release targets users who need a local, high-context, multimodal Qwen3.8 variant that prioritizes direct answers and throughput over refusal-based safety preambles. The package combines per-quant GGUF builds, a BF16 vision projector for image/video inputs, and an optional FastMTP draft sidecar that speeds speculative MTP decoding while leaving the verified target model unchanged.

What Sets It Apart
  • Aggressive uncensoring profile: configured to produce direct answers with minimal preamble and no built-in refusal behavior, useful where compliance prompts would otherwise block useful completions.
  • HauhauCS FastMTP sidecar: a 32K draft profile that can accelerate document and reasoning throughput (benchmarked up to ~3.02x document TG and ~1.93x reasoning TG versus MTP-disabled runs), while the full target still verifies each token.
  • K_P quant family and variants: per-quant "Perfect" (K_P) quantizations selectively preserve critical tensors to raise quality by ~1–2 quant levels at modest size overhead; many quant sizes offered to fit different GPU memory targets.
  • Native long-context and multimodal support: the underlying Qwen3.8 architecture supports a very large native context (262,144 tokens) and includes a vision projector for BF16 image/video inputs.
Who It's For and Trade-offs

Great fit if you need a local multimodal Qwen3.8 replica with long-context, configurable quant/size trade-offs, and optional speculative MTP acceleration for higher serving throughput. It is practical for benchmarking, research, and inference workloads where direct answers and throughput matter more than built-in refusal constraints. Look elsewhere if you require a safety-first/default-refusal model, strict moderation by default, or if you cannot run patched/runtime setups (FastMTP requires a patched llama.cpp runtime to use the sidecar). Also plan VRAM and context sizing carefully: large native context and K/V precision raise memory needs.

More Items

Hugging Face
AI Model2026

Open-weight 309B Mixture-of-Experts causal LLM with 15.5B active parameters and a native 1M-token context for coding and AI R&D. Combines Sliding-Window Attention and DeepSeek Sparse Attention (no full-attention layers), supports FP8 inference; weights under MIT license.

Hugging Face
AI Video2026

Turns a single photo into a geometry-consistent, frozen-time 360° camera orbit that returns to the exact start frame. Implemented as a LoRA for MiniMax‑H3 FL2VA — use identical first+last keyframes to produce seamless orbit clips; trained on a small human-centric square orbit dataset, so results are domain-limited.

Hugging Face
AI Audio2026

Transcribes English speech into punctuated, capitalized text — a 164 MB quantized ASR model that averages 5.21% WER across seven Open ASR Leaderboard sets. Optimized for on-device and CPU/GPU inference, with fast runtimes on Apple M5 and Docker/GPU support.