Why this matters
Small, focused decision models offer a practical alternative to brittle rule-based routers: instead of enumerating keywords and maintaining branch logic, one compact model can read context and the meanings of candidate answers together and score them directly. This flips routing from handcrafted rules to a single learned comparator that can be reused across workflows and languages.
Key Capabilities
- Decision-first interface: accepts a
state, aquestion, and 2–20options, then returns the winning option plus full softmax probabilities. That single interface covers categoricalchoice, orderedscore, and Booleannoulquestion types. - Lightweight multilingual foundation: adapts JHU CLSP’s mmBERT-small into a 144.3M-parameter checkpoint with a decision head, retaining a tokenizer and encoder suited to many languages and scenarios.
- Practical runtime profile: FP32 weights (~550.5 MiB) and a Python runtime that runs on CPU (CUDA optional for BF16 GPUs). The runtime supports up to 8,192 combined tokens in inference and a native 2–20 option limit; larger lists require hierarchical grouping.
- Measured behavior, not claims: benchmark runs show clear strengths and limits — e.g., 94/100 on an AG News pilot, 86/100 on an emotion pilot, 71.50% macro accuracy on MASSIVE scenario classification (52 locales), and ~73% on a typed-decision suite. Long, similar label lists (Banking77 pilot) reveal a concrete weakness.
Suitable users and tradeoffs
Great fit if you need a compact model to replace or augment rule-based routing and classification pipelines that accept explicit candidate answers: multilingual scenarios, CPU deployments, and systems that benefit from soft probabilities rather than hard rules. It’s suitable for prototyping routing, intent classification with closed label sets, and evaluation of decision workflows.
Look elsewhere if your task requires open-ended generation, long multi-step reasoning, or supplying missing factual information — Julia compares the answers you provide and is not designed to invent or chain long calculations. Also evaluate accuracy on your exact label wording and domain: grouping/narrowing for large choice lists can drop the correct label, and benchmark performance can differ across hardware and encoding settings.
Where it fits
Compared with general-purpose LLMs, this model is purpose-built for finite-choice decisions rather than free-form text generation. Compared with rule-based routers, it reduces maintenance overhead by scoring candidate meanings directly, but it requires careful option wording and validation on your workflow. The artifact is a runnable checkpoint + Python runtime (Apache‑2.0), not the private training pipeline.