The core insight: a behavioural "refusal" direction in Qwen3.8-Flash-Next can be removed as a rank‑1 linear edit applied in-place to an already-quantized GGUF, and that edit can be reproduced by copying bytes (tensor transplant) without re-running GSQ/RCO. That lets you test model alignment and safety trade-offs while keeping the exact learned quantization layout and scales.
Key Capabilities
- Preserves GSQ-RCO quantization: the build keeps every GSQ-learned value and block scale intact and restores the upstream per-tensor quantization types exactly, so downstream runtime behaviour is comparable except for the ablation.
- Minimal, verifiable edit: 144 "write-to-residual-stream" projection tensors across 48 layers are replaced byte-for-byte (135 by direct transplant, 9 expert tensors re-encoded with upstream GSQ block scales). Per-tensor blake2b digests verify only the intended tensors changed.
- Low overhead and deployable: four quant tiers (IQ3_S, IQ3_XXS, IQ2_XS, Q2_0) are provided; typical shard-size increase is +0.25–0.44 GB over upstream. Standard GGUF format — loads in llama.cpp and Strata with the usual pack rebuild step.
- Multimodal, long-context runtime: supports image+text pipeline, native context up to 262,144 tokens, and MoE expert routing preserved.
Who it's for & trade-offs
Great fit if you need a reproducible, low-cost subject for alignment research, red-teaming, or studies of quantization + behavioural edits: the transplant method lets you compare pre- and post-ablation while keeping the exact GSQ/RCO layout. Look elsewhere if you require a production-safe model or high-precision weights — this release intentionally removes refusal guards and provides no safety layer, and it does not include BF16/high-precision training artifacts. Reproducing or deploying the pack requires rebuilding the native pack (tensor offsets change) and applying your own moderation and access controls.