Why this matters
Swift 1.5 GSQ-RCO packs a full 27B-class model into deployable 8–12 GB GGUF files by reusing ISTA's per-tensor GSQ-RCO allocations and applying Swift-specific refinement. That makes a near-production-quality Swift 1.5 variant usable in lightweight llama.cpp runtimes while preserving distributional similarity to the BF16 model as measured by KLD across prose, code, math and multilingual text.
Key Capabilities
- Compact, mixed-precision GGUF tiers: four published tiers (IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S) with standard sizes from 8.42 GB to 11.77 GB and matching optional -mtp packages that add the MTP head. This lets you choose a size/quality tradeoff for constrained environments.
- Per-tensor GSQ + RCO workflow: reuses ISTA-DASLab's RCO allocation to assign quant types per tensor and applies GSQ refinement with a Swift V1MIX importance matrix, reducing quantization error compared to naive low-bit quants.
- Measured fidelity: held-out KLD evaluations against Swift 1.5 BF16 are reported across C4 prose, CodeParrot code, GSM8K math, and multilingual text; all refined files improve KLD over their matched starting quantizations.
- Runtime compatibility: released as standard GGUFs intended to run in llama.cpp-compatible runtimes; example invocation and context guidance are provided for deployment.
Who it's for — tradeoffs and when to use it
Great fit if you need to run a 27B-class Swift 1.5 derivative on CPU or constrained GPU environments where GGUF/llama.cpp is the runtime target, and you want per-tensor refined quantizations with published KLD fidelity metrics. Choose smaller IQ2_XS/IQ2_S tiers for minimal disk/memory footprint or IQ3 variants for higher fidelity.
Look elsewhere if you require the original BF16 weights for unconstrained training, a validated multimodal vision projector (this release does not include a verified vision projector), or if you need independent task-accuracy benchmarks at very long contexts — the reported KLD tests use 512-token contexts and distributional divergence, not end-to-end task scores.
Methods and provenance
The release reuses ISTA-DASLab's GSQ and RCO tooling for per-tensor allocation and applies Swift-specific GSQ refinement using a Swift importance matrix. The adapted weights are distributed under the Swift Open License v1.0; original Qwen components retain their Apache 2.0 terms. The package documents exact file identities, tensor allocations, and SHA256 sums for reproducibility.