Small changes in per‑token fidelity compound across long agent trajectories; OrcaSAQ2 targets that intersection by asking not just “how small can the checkpoint be?” but “what behavior survives aggressive quantization?” The result is a 12.3 GB checkpoint derived from Qwen3.8‑27B that aims to preserve next‑token distributions and downstream agent reliability while fitting practical single‑GPU memory envelopes.
Key Capabilities
- High‑fidelity compression: reduces a 54 GB BF16 checkpoint to 12.3 GB (≈77.2% storage reduction) while reporting only +0.02% perplexity and 93.2% token‑level Top‑1 agreement against BF16 on standard fidelity tests.
- Sensitivity‑aware mixed‑precision quantization: allocates bits nonuniformly so more sensitive parameters keep higher precision; decoder averages ~3.21 bits per weight.
- Agent and production features: supports a 262,144‑token architectural context, thinking mode, function/tool calling, MTP speculative decoding, and is packaged for vLLM serving (single‑stream throughput reported up to 90.1 tok/s with MTP on under a 15.7 GiB GPU cap).
- Practical single‑GPU deployment: checkpoint size and vLLM integration target 16 GB‑class GPUs for interactive agent, coding, terminal and long‑horizon tasks.
Who it's for and trade‑offs
Great fit if you need a deployable, near‑BF16 27B model for multi‑step agents (coding agents, terminal/browser agents, repository‑scale workflows) and must fit a single 16 GB GPU or similar inference footprint. It provides measurable long‑horizon evaluation points (SWE‑bench, Terminal‑Bench) as public references.
Look elsewhere if you require exact bit‑perfect reproduction of BF16 behavior (OrcaSAQ2 reports 93.2% top‑1 agreement, not 100%), image/vision inputs (the vision tower is not included), or if you need full transparency of the quantization internals (methodology, calibration and packing are proprietary and undisclosed). Serving depends on the OrcaSAQ2 vLLM integration and real usable context depends on KV cache, MTP and GPU memory headroom, so benchmark for your workload.