AIAny
AI Model2026
Icon for item

d1-3B

Maps multimodal inputs (text + images) to structured decisions (yes/no, choice, or scored rubric) in a single forward pass and returns calibrated probabilities. 3.1B parameters, long context (32,768 tokens), optimized for low-latency edge inference; not a text-generation/chat model.

Introduction

d1-3B reframes some common production needs: instead of generating free-form text, it emits a calibrated decision distribution about a presented state (text, JSON, images, or a mix) in one forward pass. That single-pass, probability-first design makes it suitable where a deterministic, auditable decision or routing choice is the desired output rather than narrative or explanation.

Key Capabilities
  • Structured decisions, not text: answers three question types—noul (binary probability), choice (named-option distribution + confidence), and score (ordered rubric with expected level and confidence). This makes integration into rule systems, routing, moderation, and agent guardrails straightforward.
  • Multimodal, long-context backbone: ~3.12B parameters with a vision encoder and 32,768-token context window, so it can reason over lengthy states and images together for a single decision.
  • Calibrated probabilities and speed: returns probabilities (no token generation) and is measured for edge and GPU: e.g., single-question latencies under 10 ms on high-end GPUs and under 50 ms on many embedded devices, enabling real-time pipelines.
  • Designed for decision pipelines: common uses include intent/topic classification, triage/routing, moderation gates, reranking, LLM-as-judge scoring, and visual inspections; it intentionally avoids free-form generation.
Who it's for and tradeoffs

Great fit if you need deterministic, auditable decisions from multimodal inputs with tight latency or edge-resource constraints — e.g., automated triage, moderation filters, model routing, or tool-call approval in agents. The probability outputs simplify downstream thresholds and monitoring.

Look elsewhere if your application requires open-ended generation, detailed explanations, or conversational AI: this model does not produce text responses and is not optimized for creative or explanatory outputs. Also, while Decision Index and benchmark numbers are strong for its size, accuracy varies by task and you should validate calibration and label definitions for your domain before relying on automated gating.

More Items

Hugging Face
AI Model2026

Performs a byte-level transplant of 144 tensors in an already-quantized GSQ-RCO Qwen3.8-Flash-Next to ablate the model's refusal direction while preserving GSQ-learned scales and the upstream per-tensor type assignment; multimodal, 262K context. Intended for local inference, red-teaming and quantization research; no retraining or built-in safety.

Hugging Face
AI Model2026

Runs a pruned, NVFP4-quantized GLM-5.3-Flash variant tuned for Blackwell GPUs: 224 routed experts per layer and ~141 GiB of weights. Retains the multimodal vision tower, activates 18B params/token, supports vLLM and optional MTP speculative decoding; fits 2× DGX Spark or a ≥180 GB B200.

Hugging Face
AI Model2026

Returns calibrated probabilities for yes/no and multi-choice decisions using a two-stage System 1 (fast classifier) and System 2 (Gemma‑4 reasoning) pipeline; this NVFP4 release quantizes the 3,840 routed MoE experts to 4-bit so the model fits ≈17–18 GB of GPU memory while other weights remain bf16.