Run decision models locally.
Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware.
An independent open-source project, not affiliated with Ollama or TypeSafe.
ollaya run winnow:e4b \"Third time this year you'vedouble-charged me. Refund ittoday or I'm cancelling andmoving to a competitor."| Question | Answer | Probability |
|---|---|---|
| intent | refund | 0.91 |
| is_urgent | yes | 0.92 |
| frustration | 2.89 / 3 very angry | 0.86 |
| refund_requested | yes | 0.99 |
| churn_risk | yes | 0.99 |
Fast and accurate
Close to Jev's accuracy, in under 100 ms.
A decision model answers in a single forward pass, with no token-by-token generation. On an RTX 4090, winnow:e4b answers a five-question request in 89 ms end to end, and scores 0.722 on typed decisions against 0.738 for TypeSafe's hosted Jev. Smaller models such as laya answer in about 10 ms, and run well on a CPU.
89ms
winnow:e4b on Ollaya
RTX 4090, five questions · 0.722 accuracy
236–276ms
TypeSafe Jev
Hosted API, median request · 0.738 accuracy
- TypeSafe Jevhosted API0.738236–276 ms
- winnow:e4b0.72289 ms
- kev:9b0.722498 ms
- winnow:12b0.702131 ms
- decider:4b0.680520 ms
- kev:4b0.669354 ms
- jevk5:4b0.625105 ms
- decider:2b0.591190 ms
- nli0.54820 ms
- decider:0.8b0.506155 ms
- gliclass0.47715 ms
- kev:0.8b0.460128 ms
- von0.44723 ms
- laya:en0.36110 ms
- clm:8bquestions cached0.357149 ms
Accuracy: the typed-decisions test split (400 states, 2,000 questions), argmax against the majority label, measured by Ollaya for each model; Jev's from Winnow's benchmark report on the same questions. laya:typed-decisions scores 0.766 but was fine-tuned on this dataset, so it is left out. clm caches questions and options: its latency is a new message whose questions are cached. Latency: median of a five-question request through the HTTP API on an NVIDIA RTX 4090; Jev: median request of the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), which includes the network. Setups differ, so read the latencies as an order-of-magnitude comparison.
Drop-in compatible
Speaks TypeSafe's API.
Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.
Request
# Point the TypeSafe SDK at Ollaya
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value works
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
# …or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d '{
"model": "winnow:e4b",
"state": "Can I get an invoice for last month?",
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"invoice": "Needs an invoice or receipt",
"refund": "Wants money back",
"other": "Anything else"
}
}
}
}'Response
{
"model": "winnow:e4b",
"answers": {
"intent": {
"type": "choice",
"choice": "invoice",
"confidence": 0.9801,
"probabilities": {
"invoice": 0.9868,
"refund": 0.0026,
"other": 0.0106
}
}
},
"usage": {
"input_tokens": 120,
"output_tokens": 0
}
}Open models
Open weights, ready to pull.
Pick by what you need: winnow:e4b balances accuracy and speed best, laya is the fastest and runs well on a CPU, kev and decider scale up to 9B and 4B, von reads up to 8,192 tokens, and qwen3guard screens text for safety. The models page shows each one’s accuracy and speed.
- winnowDecision models by EldanRing, fine-tuned from Google's Gemma 4 and published as GGUF. Winnow reads the answer labels' logits after its own prompt; Ollaya runs the author's file on llama.cpp, on NVIDIA GPUs, Apple silicon or the CPU.7.5b · 12b
- layaOpen decision models from Convai Innovations. Typed, calibrated answers to choice, score and yes/no questions in a single forward pass, in English and 100+ languages.322m · 421m
- deciderDecoder decision models by Mapika on Qwen3.5: the answer is read from option-letter logits in one forward pass. decider:4b scores 0.680 on typed decisions, decider:2b 0.591.0.75b · 1.9b · 4.2b
- kevDecision models by Jared Palmer: a LoRA on a Qwen3.5 base plus a pointer head that scores every option at its own span, in one forward pass per question. Calibrated with Kev's own temperature.0.76b · 4.2b · 7.9b
- nliZero-shot classifiers by Moritz Laurer: every option becomes a hypothesis scored for entailment. The most accurate encoder model on typed decisions in our tests.396m · 435m
- gliclassInstruction-following zero-shot classifier by Knowledgator: all options of a question are scored in one pass, so cost barely grows with the number of options.439m
- qwen3guardSafety guard by the Qwen team: is a text safe, controversial or unsafe, and which unsafe category? It answers its own built-in questions, in 119 languages, in one forward pass.0.6b
- decisionDecision models by the vLLM Semantic Router contributors: a fully fine-tuned Qwen3.5 backbone plus an endpoint head that scores every option at its own last token against the question, in one forward pass per question. 16k-token rows.0.75b
Your data stays yours
Private by default.
Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.
Local
Runs on your machine with ONNX Runtime, on the CPU or an NVIDIA GPU. The server listens on 127.0.0.1 by default.
Open weights
Weights come from their authors’ Hugging Face repositories, pinned to a commit and checked against sha256. Ollaya never re-hosts them, and the runtime is Apache-2.0.
No per-token fees
Run as many decisions as your hardware can handle. No metering and no API bill.
Calibrated
Probabilities you can put thresholds on. Each model ships its own calibration, and a Modelfile refits it on your labelled data.
Platforms
Runs where you work.
A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, Windows, WSL 2 or Docker takes a request down to milliseconds.
| Platform | Desktop app | Command line | GPU |
|---|---|---|---|
| macOSApple silicon, macOS 14+ | Desktop appMenu bar app.dmg | Command lineInstall script | GPUApple GPULaya and NLI on MLX |
| Windows10 and 11, x64 | Desktop appDesktop app.exe or .msi | Command linePowerShell script | GPUNVIDIA, CUDA 13 or 12Command line |
| Linuxx86-64 | Desktop appDesktop appAppImage, .deb, .rpm | Command lineInstall scriptsystemd service | GPUNVIDIA, CUDA 13 or 12 |
| LinuxARM64 | Desktop appNot available | Command lineInstall scriptsystemd service | GPUCPU only |
| WSL 2Linux on Windows | Desktop appNot available | Command lineInstall scriptSame as Linux | GPUNVIDIA, CUDA 13 or 12 |
| Dockeramd64 and arm64 | Desktop appNot available | Command lineImage on GHCR | GPUNVIDIA, CUDA 13 or 12:cuda and :cuda12, amd64 |
NVIDIA GPUs need driver R525 or newer; the install scripts fetch the CUDA libraries only when they find one. On a Mac, laya and nli run on the Apple GPU through MLX; other models, AMD and Intel GPUs, and the Windows and Linux desktop apps use the CPU.
Get up and running in minutes.
One binary, one command: ollaya run winnow:e4b.
macOS, Windows, Linux and Docker · Apache-2.0 · GitHub