Most pipelines expect a model to generate text and then parse it; Clef flips that flow by producing strictly typed, probability-weighted decisions directly from the input state. That makes it practical to route tickets, apply policy rubrics, or gate agent actions with a single low-latency call and no output parsing.
Key Capabilities
- Joint, schema-bound scoring: a small transformer head attends across the backbone's final hidden states and scores every allowed option for every question in parallel, returning one logit per option (softmax per question yields probabilities).
- Multimodal input and large context: accepts text, JSON, images, and video frames with a 64k token context window, so you can include long documents or visual evidence alongside structured state.
- Typesafe outputs and API compatibility: supports
noul(boolean probability),choice(named options with per-option probabilities and confidence), andscore(ordered rubric with expected score). Fully compatible with Jev / SystemOne request/response formats. - Engineering tradeoffs: Clef is a 27B model post-trained from Qwen/Qwen3.8-27B (vision encoder included). Clef-flash (9B) is offered when latency is critical.
Practical notes and limits
- Inference returns probabilities, not free-form text, eliminating parsing errors but also meaning Clef isn't suitable where natural-language explanations are required.
- Image/video support: embedded images (client-side) may be included in records; Workers AI docs limit embedded PNG/JPEG/WebP images to up to 4 images (size and decoded limits apply). Text+multimodal records can be batched together.
- Deployment requirements: tested with PyTorch 2.11 and transformers 5.10.2 on an H200; image/video processing requires Pillow. Model files are provided as sharded safetensors; license is Apache-2.0.
Who it's for, and when to choose something else
- Great fit if you need deterministic, structured decisions from mixed inputs (e.g., automated ticket routing, trust & safety rubrics, invoice/workflow routing, agent guardrails) and value calibrated probabilities over free text.
- Choose Clef-flash or smaller classification models when sub-100ms latency is essential or when GPU/memory budgets prohibit a 27B model. Avoid Clef if you require fluent, explainable free-text outputs or lightweight on-device inference.