AIAny
AI Model2026
Icon for item

OrcaSAQ2 27B

Compresses Qwen3.8-27B into a 12.3 GB sensitivity-aware mixed-precision quantized checkpoint for long-horizon agent workloads; preserves BF16 fidelity (+0.02% PPL, 93.2% token Top‑1 agreement), supports 262K context and vLLM serving, text-only and Apache‑2.0 licensed.

Introduction

Small changes in per‑token fidelity compound across long agent trajectories; OrcaSAQ2 targets that intersection by asking not just “how small can the checkpoint be?” but “what behavior survives aggressive quantization?” The result is a 12.3 GB checkpoint derived from Qwen3.8‑27B that aims to preserve next‑token distributions and downstream agent reliability while fitting practical single‑GPU memory envelopes.

Key Capabilities
  • High‑fidelity compression: reduces a 54 GB BF16 checkpoint to 12.3 GB (≈77.2% storage reduction) while reporting only +0.02% perplexity and 93.2% token‑level Top‑1 agreement against BF16 on standard fidelity tests.
  • Sensitivity‑aware mixed‑precision quantization: allocates bits nonuniformly so more sensitive parameters keep higher precision; decoder averages ~3.21 bits per weight.
  • Agent and production features: supports a 262,144‑token architectural context, thinking mode, function/tool calling, MTP speculative decoding, and is packaged for vLLM serving (single‑stream throughput reported up to 90.1 tok/s with MTP on under a 15.7 GiB GPU cap).
  • Practical single‑GPU deployment: checkpoint size and vLLM integration target 16 GB‑class GPUs for interactive agent, coding, terminal and long‑horizon tasks.
Who it's for and trade‑offs

Great fit if you need a deployable, near‑BF16 27B model for multi‑step agents (coding agents, terminal/browser agents, repository‑scale workflows) and must fit a single 16 GB GPU or similar inference footprint. It provides measurable long‑horizon evaluation points (SWE‑bench, Terminal‑Bench) as public references.

Look elsewhere if you require exact bit‑perfect reproduction of BF16 behavior (OrcaSAQ2 reports 93.2% top‑1 agreement, not 100%), image/vision inputs (the vision tower is not included), or if you need full transparency of the quantization internals (methodology, calibration and packing are proprietary and undisclosed). Serving depends on the OrcaSAQ2 vLLM integration and real usable context depends on KV cache, MTP and GPU memory headroom, so benchmark for your workload.

Information

  • Websitehuggingface.co
  • OrganizationsOrcaRouter, Continuum AI (Continuum-AI-Corp)
  • Published date2026/09/24

Categories

More Items

Hugging Face
AI Model2026

Turns a context, a question, and 2–20 candidate answers into a single chosen option for classification, routing, ordered scores, and Boolean decisions. Built on mmBERT-small with a 144.3M-parameter decision head, runs on CPU with up to 8,192 combined tokens; FP32 weights occupy 550.5 MiB and are Apache‑2.0 licensed.

Hugging Face
AI Model2026

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.

Hugging Face
AI Model2026

Parses digital and camera-captured documents into structured outputs (text, layout, tables, formulas, figures) using a lightweight (~1.2B) open-source vision-language model. Uses geometry-aware modeling, multi-node consensus pseudo-labeling, and content-structure decoupling to handle warped, photographed, and digital pages.