Models
Open decision models you can run locally with Ollaya.
18 models
winnow
Decision models by EldanRing, fine-tuned from Google's Gemma 4 and published as GGUF. Winnow reads the answer labels' logits after its own prompt; Ollaya runs the author's file on llama.cpp, on NVIDIA GPUs, Apple silicon or the CPU.
decisionmultilingualfine-tunedgguf7.5b12b5 TagsUpdated 0.722 typed-decisions89 ms on RTX 4090
laya
Open decision models from Convai Innovations. Typed, calibrated answers to choice, score and yes/no questions in a single forward pass, in English and 100+ languages.
multilingualrouterguardrailsfine-tuned322m421m10 TagsUpdated 0.361 typed-decisions9.6 ms on RTX 4090
decider
Decoder decision models by Mapika on Qwen3.5: the answer is read from option-letter logits in one forward pass. decider:4b scores 0.680 on typed decisions, decider:2b 0.591.
decisionlong-contextvision0.75b1.9b2.2b4.2b5 TagsUpdated 0.680 typed-decisions520 ms on RTX 4090
kev
Decision models by Jared Palmer: a LoRA on a Qwen3.5 base plus a pointer head that scores every option at its own span, in one forward pass per question. Calibrated with Kev's own temperature.
decision0.76b4.2b7.9b4 TagsUpdated 0.669 typed-decisions354 ms on RTX 4090
nli
Zero-shot classifiers by Moritz Laurer: every option becomes a hypothesis scored for entailment. The most accurate encoder model on typed decisions in our tests.
zero-shot396m435m3 TagsUpdated 0.548 typed-decisions20.4 ms on RTX 4090
gliclass
Instruction-following zero-shot classifier by Knowledgator: all options of a question are scored in one pass, so cost barely grows with the number of options.
zero-shot439m2 TagsUpdated 0.477 typed-decisions14.7 ms on RTX 4090
qwen3guard
Safety guard by the Qwen team: is a text safe, controversial or unsafe, and which unsafe category? It answers its own built-in questions, in 119 languages, in one forward pass.
guardrailsmultilingual0.6b2 TagsUpdated 37 ms on RTX 4090
decision
Decision models by the vLLM Semantic Router contributors: a fully fine-tuned Qwen3.5 backbone plus an endpoint head that scores every option at its own last token against the question, in one forward pass per question. 16k-token rows.
decisionlong-context0.75b2 TagsUpdated 217 ms on RTX 4090
von
Decision model by Victor Hugo Panisa on ModernBERT-large: every option is scored at its own marker, all options of a question in one pass, with an input-conditioned calibration. 8k-token context.
decisionlong-context395m2 TagsUpdated 0.447 typed-decisions23 ms on RTX 4090
jevk5
Decision model by alibiserikbay, fine-tuned from Qwen3.5-4B and published as GGUF. JevK5 reads the answer letters' logits after its own JSON prompt; Ollaya runs the author's file on llama.cpp, on NVIDIA GPUs, Apple silicon or the CPU.
decisionfine-tunedgguf4b2 TagsUpdated 0.625 typed-decisions105 ms on RTX 4090
clm
Contrastive decision model by Contrastive-LM: the Qwen3-8B encoder embeds the state and every option, and two trained heads pick the option closest to the state. Repeated questions and options are cached.
decisionembedding8.2b2 TagsUpdated 0.357 typed-decisions149 ms on RTX 4090
nimble
Decision model by Bespoke Labs: a LoRA on Qwen3.5-9B trained on contrastive pairs. Nimble reads the whole request as a JSON schema and scores each option by the next-token logit of its code; Ollaya applies the author's temperature, so its probabilities are calibrated. Up to 255 options.
decisionfine-tuned9b2 TagsUpdated 0.665 typed-decisions2297 ms on RTX 4090
jeb
Decision models by Jason Brashear and AINode: LoRAs merged into Qwen3.5 and Qwen3.8, published as GGUF. Jeb reads the option labels' logits after AINode's own decision prompt, with a fitted temperature per question type; Ollaya runs the authors' files on llama.cpp.
decisionfine-tunedgguf4b9b27b4 TagsUpdated 124 ms on RTX 4090
cygnet
Frozen Gemma 4 12B IT with blockbrain-ai's Cygnet prompt: the options as letters, one answer slot, and one calibration temperature. No fine-tuning; Ollaya runs ggml-org's Q8_0 GGUF on llama.cpp.
decisionmultilingualgguf12b2 TagsUpdated 0.683 typed-decisions202 ms on RTX 4090
jeeves
Decision model by PostHog: Qwen3.5-9B with a LoRA merged in and a pointer head that scores every option at its own marker. Ollaya runs it without its reasoning chain, in one forward pass per question, calibrated with the authors' temperature.
decisionfine-tuned9b2 TagsUpdated 0.680 typed-decisions838 ms on RTX 4090
clef
Decision models by Cloudflare. Clef-Flash is Qwen3.5-9B, fully post-trained, with a joint schema head that scores every option of every question together, in one forward pass per request. Its probabilities come straight from the head. Ollaya runs its text path.
decisionfine-tuned9b2 TagsUpdated 0.703 typed-decisions532 ms on RTX 4090
snap
logitlab's snap1-2b: MiniCPM5-2B fine-tuned to read one option letter after the prompt of emnlmn's snap engine. Ollaya runs the author's Q8_0 GGUF on llama.cpp with snap's own prompt, on NVIDIA GPUs, Apple silicon or the CPU.
decisionfine-tunedgguf2b2 TagsUpdated 0.648 typed-decisions68 ms on RTX 4090
decima
A. M. Madani's Decima: multilingual encoders (mmBERT-base, multilingual-e5-small) with a late-interaction scorer that reads every option against the state, so option order never changes the answer, and an ordinal head for scores. A general model, one for coding-agent decisions, and a small one that is the fastest on a CPU.
decisionmultilingualfine-tuned122m321m4 TagsUpdated 0.495 typed-decisions15.1 ms on RTX 4090
Numbers are for each model's default tag. Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states; the labels have low annotator agreement, so compare models with each other rather than reading the numbers as absolutes. Latency is the median five-question request with a short state, end to end through the HTTP API on an RTX 4090 (qwen3guard: its four built-in questions). Larger models are more accurate and slower; see each model's page for how speed grows with the length of the state.
No models found.