AIAny
AI Model2026
Icon for item

Swift 1.5 Qwen3.8-27B · GSQ-RCO

Provides compact mixed-precision GGUF quantizations of UkisAI's Swift 1.5 (derived from Qwen3.8-27B), using ISTA GSQ-RCO per-tensor allocations with Swift-specific refinement. Offers multiple 8–12 GB tiers, optional MTP heads, and KLD evaluation against the Swift BF16 baseline.

Introduction

Why this matters

Swift 1.5 GSQ-RCO packs a full 27B-class model into deployable 8–12 GB GGUF files by reusing ISTA's per-tensor GSQ-RCO allocations and applying Swift-specific refinement. That makes a near-production-quality Swift 1.5 variant usable in lightweight llama.cpp runtimes while preserving distributional similarity to the BF16 model as measured by KLD across prose, code, math and multilingual text.

Key Capabilities
  • Compact, mixed-precision GGUF tiers: four published tiers (IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S) with standard sizes from 8.42 GB to 11.77 GB and matching optional -mtp packages that add the MTP head. This lets you choose a size/quality tradeoff for constrained environments.
  • Per-tensor GSQ + RCO workflow: reuses ISTA-DASLab's RCO allocation to assign quant types per tensor and applies GSQ refinement with a Swift V1MIX importance matrix, reducing quantization error compared to naive low-bit quants.
  • Measured fidelity: held-out KLD evaluations against Swift 1.5 BF16 are reported across C4 prose, CodeParrot code, GSM8K math, and multilingual text; all refined files improve KLD over their matched starting quantizations.
  • Runtime compatibility: released as standard GGUFs intended to run in llama.cpp-compatible runtimes; example invocation and context guidance are provided for deployment.
Who it's for — tradeoffs and when to use it

Great fit if you need to run a 27B-class Swift 1.5 derivative on CPU or constrained GPU environments where GGUF/llama.cpp is the runtime target, and you want per-tensor refined quantizations with published KLD fidelity metrics. Choose smaller IQ2_XS/IQ2_S tiers for minimal disk/memory footprint or IQ3 variants for higher fidelity.

Look elsewhere if you require the original BF16 weights for unconstrained training, a validated multimodal vision projector (this release does not include a verified vision projector), or if you need independent task-accuracy benchmarks at very long contexts — the reported KLD tests use 512-token contexts and distributional divergence, not end-to-end task scores.

Methods and provenance

The release reuses ISTA-DASLab's GSQ and RCO tooling for per-tensor allocation and applies Swift-specific GSQ refinement using a Swift importance matrix. The adapted weights are distributed under the Swift Open License v1.0; original Qwen components retain their Apache 2.0 terms. The package documents exact file identities, tensor allocations, and SHA256 sums for reproducibility.

Information

  • Websitehuggingface.co
  • OrganizationsUkisAI, Deep Algorithms and Systems Lab, Institute of Science and Technology Austria (ISTA-DASLab), Alibaba Cloud (Qwen team)
  • Published date2026/09/24

Categories

More Items

Hugging Face
AI Model2026

Capability-targeted compression of Qwen3.8-Flash-Next: half the experts are removed and remaining weights quantized to 3.5 bpw, producing a 58.4 GB GGUF (29.6 GB resident) that preserves coding and multimodal ability while trading off other domains.

Hugging Face
AI Model2026

Provides a high-fidelity mixed-precision (≈3-bit) GGUF quant of Qwen3.8-27B tailored for long-horizon agent workloads and cyber-focused red-teaming. Preserves reasoning, thinking mode, MTP speculative decoding and vision (via a separate mmproj); released as an uncensored/abliterated research build under Apache-2.0.

Hugging Face
AI Model2026

Provides official pretrained VisionHOPE visual-backbone checkpoints for ImageNet classification, COCO object detection & instance segmentation, and ADE20K semantic segmentation. Includes hierarchical Tiny/Small/Base models with PyTorch-compatible downloadable weights.