AIAny
AI Model2026
Icon for item

GLiNER2.5-Decide

Performs schema-driven, low-latency classification and structured decision-making over English text. Supports multi-head scoring, constrained joint decoding with confidence/feasibility metadata, span extraction, and local CPU/GPU deployment via the gliner2 runtime.

Introduction

Production software often needs deterministic, constrained decisions from short text (routing, triage, policy tags) rather than open-ended generation. GLiNER2.5-Decide is a compact encoder-first decision model fine-tuned to return valid, jointly consistent answers to typed questions and label sets in a single forward pass, trading large-model reasoning for speed, predictability, and deployability.

Key Capabilities
  • Schema-driven decisions: accepts a declarative schema of typed questions (single-label, multi-label, ordinal) and permitted answers; label descriptions and rules can be included so the model reasons about label compatibility rather than free-form text.
  • Constrained joint decoding: computes compatibility scores for every candidate answer and finds the highest-scoring assignment that satisfies declared constraints, returning selected answers plus probabilities, per-head confidence, and feasibility metadata.
  • Multi-task, low-latency operation: one forward call can score multiple heads (intent, routing, sentiment, urgency, policy, etc.) and multi-label sets without generating tokens; reported p50 latencies are low (tens of ms on GPU, hundreds of ms on large CPU machines for short documents).
  • Practical deployment: 340M-parameter encoder (DeBERTa-v3-large backbone), runs locally on CPU or GPU via the gliner2 library, Apache-2.0 license, and supports fine-tuning and LoRA.
  • Benchmarks & behavior: leads an internal 17-dataset "fast-decisions" benchmark (~60.1% exact-match average) and emphasizes operational decision accuracy over open-ended reasoning or explanations.
Who it's for & trade-offs

Great fit if you need deterministic, auditable routing/triage/moderation/triage decisions from text at low latency and want a compact model that runs in production without relying on large decoder LLMs. Look elsewhere if you require generative explanations, chain-of-thought reasoning, or broad open-domain QA—GLiNER2.5-Decide is explicitly not a general-purpose reasoning model and was optimized for structured, operational decision-making rather than benchmark-spanning language understanding.

Information

  • Websitehuggingface.co
  • OrganizationsFastino AI
  • AuthorsUrchade Zaratiana, Gil Pasternak, Oliver Boyd, George Hurn-Maloney, Ash Lewis
  • Published date2026/09/23

Categories

More Items

Hugging Face
AI Model2026

Parses digital and camera-captured documents into structured outputs (text, layout, tables, formulas, figures) using a lightweight (~1.2B) open-source vision-language model. Uses geometry-aware modeling, multi-node consensus pseudo-labeling, and content-structure decoupling to handle warped, photographed, and digital pages.

Hugging Face
AI Model2026

Scans long documents rendered as compressed page-images, locates relevant pages, and selectively expands only those pages to full text for question answering; built on Qwen3.5-9B, supports 5x/10x/15x compression and is released under Apple’s research-only model license.

Hugging Face
AI Model2026

Scores candidate actions against a textual state using contrastive state/action embeddings for very fast zero-shot ranking and typed decision-making. Built as two small projection heads on frozen Qwen3-8B; fine-tunable as a verifier for agentic benchmarks.