Most real-world agents need many quick, well-calibrated decisions rather than long-form generation. JEV-27B-VL separates a fast System 1 decision head (LoRA adapter) that returns calibrated probabilities in one forward pass from an untouched Qwen3.8-27B System 2 for deliberate multimodal reasoning — and extends both modes to image inputs zero-shot.
Key Capabilities
- Fast, calibrated decisions: yes/no (
noul), multi-way choice (2–256 options), and 0–5 scoring in a single forward pass, with very low calibration error and high fidelity to TypeSafe Jev distributions. - Multimodal support: decision head applied zero-shot to images (System 1) and full image-aware generation/reasoning via the Qwen3.8-27B backbone (System 2).
- Production-friendly API and scale: serves via vLLM with a dedicated POST /v1/decide endpoint, supports up to 256K-token prompts, LoRA adapters for System 1, and Apache-2.0 weights.
- Proven applied strengths: top multimodal judge on VL-RewardBench (78.3% accuracy), zero-shot short-video recommendation from covers (AUC 0.727, matching collaborative filtering), strong agent-judging and search re-ranking results.
Who it's for and trade-offs
Great fit if you need high-throughput, calibrated automated judgments over text and images (content moderation, recommendation ranking, UI action selection, agent evaluation) and want to combine these with a full reasoning LLM in the same deployment. Look elsewhere if you require fully supervised, image-trained decision heads (image decisions here are zero-shot and their calibration is less systematically measured) or if you cannot allocate GPU memory for large-context multimodal runs (full 256K context raises memory requirements). The model trades occasional extra latency for very wide context and a single unified weights+adapter deployment.
Where it fits
JEV-27B-VL is positioned between specialized closed hosted decision services and large multimodal generative models: it delivers calibrated decision APIs suitable for real-time ranking/control loops while retaining the full Qwen3.8-27B reasoning ability for escalation to System 2. Use it when you need deterministic, interpretable probabilities at scale and an integrated path to deeper multimodal explanations.