AIAny
AI Model2026
Icon for item

Clef-Flash

Converts a state and a typed-question schema into probabilistic, structured decisions from text, JSON, images, or video—returning per-option probabilities for noul/choice/score questions in a single forward pass. Post-trained from Qwen3.5-9B and compatible with Jev/SystemOne APIs; optimized for low-latency decisioning.

Introduction

Most classification or routing systems either emit free-form text to be parsed or run slow autoregressive passes. Clef-Flash takes a different approach: it directly scores every allowed option for each typed question in a single forward pass, producing calibrated probabilities you can act on immediately—making it suited to high-throughput, latency-sensitive decision pipelines.

Key Capabilities
  • Schema-first outputs: supports three typed question kinds (noul/choice/score) and returns per-option probabilities, a confidence metric, and probability-weighted scores for ordered rubrics.
  • Multimodal inputs: accepts text, JSON, images, and video frames and routes evidence from media and text to question-specific scorers via a joint schema head.
  • Lightweight, low-latency variant: a 9B post-trained Qwen3.5-9B backbone with a small joint transformer head; designed for hot-path decisions where median inference latency is substantially lower than larger autoregressive models.
  • API compatibility: follows the Jev/SystemOne request/response shape so it can replace schema-based decision endpoints without introducing free-form outputs to parse.
Who it's for and trade-offs

Great fit if you need deterministic, schema-bound decisions in real time—examples include ticket routing, automated form triage, receipt/receipt-ocr checks, and alert escalation where a probability distribution over fixed actions is required. Look elsewhere if you need free-form generation, extensive instruction-following, or the highest-precision, large-context reasoning (the larger Clef 27B model targets those cases). Clef-Flash is optimized for decisioning rather than open-ended text synthesis and has operational requirements (GPU inference stack, tested with recent torch/transformers and H200 in upstream docs).

More Items

Hugging Face
AI Model2026

Turns a structured state (text/JSON/images/video) plus a schema of typed questions into probabilistic decisions in one forward pass. Multimodal, Jev/SystemOne-compatible, 64k context window, post-trained from Qwen3.8-27B; no free-form text output.

Hugging Face
AI Model2026

Capability-targeted compression of Qwen3.8-Flash-Next: half the experts are removed and remaining weights quantized to 3.5 bpw, producing a 58.4 GB GGUF (29.6 GB resident) that preserves coding and multimodal ability while trading off other domains.

Hugging Face
AI Model2026

Provides compact mixed-precision GGUF quantizations of UkisAI's Swift 1.5 (derived from Qwen3.8-27B), using ISTA GSQ-RCO per-tensor allocations with Swift-specific refinement. Offers multiple 8–12 GB tiers, optional MTP heads, and KLD evaluation against the Swift BF16 baseline.