Most classification or routing systems either emit free-form text to be parsed or run slow autoregressive passes. Clef-Flash takes a different approach: it directly scores every allowed option for each typed question in a single forward pass, producing calibrated probabilities you can act on immediately—making it suited to high-throughput, latency-sensitive decision pipelines.
Key Capabilities
- Schema-first outputs: supports three typed question kinds (noul/choice/score) and returns per-option probabilities, a confidence metric, and probability-weighted scores for ordered rubrics.
- Multimodal inputs: accepts text, JSON, images, and video frames and routes evidence from media and text to question-specific scorers via a joint schema head.
- Lightweight, low-latency variant: a 9B post-trained Qwen3.5-9B backbone with a small joint transformer head; designed for hot-path decisions where median inference latency is substantially lower than larger autoregressive models.
- API compatibility: follows the Jev/SystemOne request/response shape so it can replace schema-based decision endpoints without introducing free-form outputs to parse.
Who it's for and trade-offs
Great fit if you need deterministic, schema-bound decisions in real time—examples include ticket routing, automated form triage, receipt/receipt-ocr checks, and alert escalation where a probability distribution over fixed actions is required. Look elsewhere if you need free-form generation, extensive instruction-following, or the highest-precision, large-context reasoning (the larger Clef 27B model targets those cases). Clef-Flash is optimized for decisioning rather than open-ended text synthesis and has operational requirements (GPU inference stack, tested with recent torch/transformers and H200 in upstream docs).