AIAny
AI Model2026
Icon for item

autotrust/GEV-26B-Decide-NVFP4

Returns calibrated probabilities for yes/no and multi-choice decisions using a two-stage System 1 (fast classifier) and System 2 (Gemma‑4 reasoning) pipeline; this NVFP4 release quantizes the 3,840 routed MoE experts to 4-bit so the model fits ≈17–18 GB of GPU memory while other weights remain bf16.

Introduction

Quantizing the routed experts to NVIDIA NVFP4 is the core insight: it reduces the memory footprint of a 26B Gemma‑4 Decide model so the same System 1 (fast calibrated decision head) and System 2 (Gemma‑4 reasoning) pipeline can run on a single 24GB-class GPU without retraining the adapter.

What Sets It Apart
  • Expert-focused NVFP4 quantization: only the 3,840 routed MLP experts (30 layers × 128 experts) are converted to a 4‑bit floating format while attention, dense MLPs, routers, vision tower and decision heads remain bf16. This targets the weights that dominate memory use and preserves numerical fidelity where it matters.
  • Much smaller footprint, same behavior: weights fall from ~49.5 GB to ~18 GB (vLLM weight memory ≈17.1 GiB), enabling deployment on 24GB GPUs and reducing GPU memory for weights by ~66% while keeping System 1’s calibration and System 2 reasoning intact.
  • One‑engine vLLM deployment and readout: the package includes vLLM-compatible artifacts and a serve.sh that exposes a /v1/decide API implementing System 1 decisions, adaptive thinking (System 1 → System 2 when confidence < threshold) and the calibrated probability readout rather than generation.
  • Decision-centric design and multimodal inputs: System 1 returns a probability distribution over 2–256 options in one forward pass (text and images supported zero‑shot); adaptive thinking folds Gemma‑4 reasoning into final probabilities when needed.
Who it’s for and trade‑offs

Great fit if you need a deployable decision model that: runs on a single 24GB GPU, produces calibrated option probabilities for automation or UI gating, or participates in low‑latency control loops (System 1 median ≈45 ms on a B200). The NVFP4 checkpoint is optimized for vLLM and keeps the System 1 adapter unchanged, so existing decision workflows port with minimal changes.

Look elsewhere if you require the absolute highest precision across every submodule (some non‑expert weights remain bf16) or if you need the full bf16 evaluation sweep used for Decision Index scoring; adaptive thinking (System 2) can be very slow on long reasoning budgets and may hurt classification tasks that System 1 was explicitly trained on. The model is English‑centric and should not be used for high‑stakes decisions without confidence gating.

Summary: the release is a practical engineering trade — keep the model’s decision behavior and reasoning stack while cutting memory by quantizing the largest weight component (routed experts), enabling real‑world deployment on smaller GPUs.

More Items

Hugging Face
AI Model2026

GGUF-packaged weights for EmbeddingGemma 2 enabling local multimodal embeddings (text, image, video, audio); supports 768/512/256/128 dimensions, BF16/FP32 inference, selective modality loading for lower memory, and is suited for semantic search and RAG.

Hugging Face
AI Model2026

Converts Turkish text to 48 kHz speech offline using a compact 5.6M DiT acoustic model plus a 3M vocoder (8.6M parameters, ~34 MB). Streams audio with very low latency (~3.86 ms first audio on an RTX 4090), supports fast batching and runs under an Apache‑2.0 license.

Hugging Face
AI Model2026

Generates unified 768‑dimensional embeddings for text (including code), images, video and audio to enable cross‑modal semantic search and retrieval. Supports task instruction prefixes, Matryoshka truncation to 128/256/512/768 dims, and modular encoders for on‑device use under an Apache‑2.0 license.