Production software often needs deterministic, constrained decisions from short text (routing, triage, policy tags) rather than open-ended generation. GLiNER2.5-Decide is a compact encoder-first decision model fine-tuned to return valid, jointly consistent answers to typed questions and label sets in a single forward pass, trading large-model reasoning for speed, predictability, and deployability.
Key Capabilities
- Schema-driven decisions: accepts a declarative schema of typed questions (single-label, multi-label, ordinal) and permitted answers; label descriptions and rules can be included so the model reasons about label compatibility rather than free-form text.
- Constrained joint decoding: computes compatibility scores for every candidate answer and finds the highest-scoring assignment that satisfies declared constraints, returning selected answers plus probabilities, per-head confidence, and feasibility metadata.
- Multi-task, low-latency operation: one forward call can score multiple heads (intent, routing, sentiment, urgency, policy, etc.) and multi-label sets without generating tokens; reported p50 latencies are low (tens of ms on GPU, hundreds of ms on large CPU machines for short documents).
- Practical deployment: 340M-parameter encoder (DeBERTa-v3-large backbone), runs locally on CPU or GPU via the gliner2 library, Apache-2.0 license, and supports fine-tuning and LoRA.
- Benchmarks & behavior: leads an internal 17-dataset "fast-decisions" benchmark (~60.1% exact-match average) and emphasizes operational decision accuracy over open-ended reasoning or explanations.
Who it's for & trade-offs
Great fit if you need deterministic, auditable routing/triage/moderation/triage decisions from text at low latency and want a compact model that runs in production without relying on large decoder LLMs. Look elsewhere if you require generative explanations, chain-of-thought reasoning, or broad open-domain QA—GLiNER2.5-Decide is explicitly not a general-purpose reasoning model and was optimized for structured, operational decision-making rather than benchmark-spanning language understanding.