Skip to content

Run decision models locally.

Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware.

An independent open-source project, not affiliated with Ollama or TypeSafe.

ollaya run winnow:e4b \
"Third time this year you've
double-charged me. Refund it
today or I'm cancelling and
moving to a competitor."
Answers returned by the model
QuestionAnswerProbability
intentrefund0.91
is_urgentyes0.92
frustration2.89 / 30.86
refund_requestedyes0.99
churn_riskyes0.99
Real output: winnow:e4b answered five questions in 87 ms on an RTX 4090.

Fast and accurate

Close to Jev's accuracy, in under 100 ms.

A decision model answers in a single forward pass, with no token-by-token generation. On an RTX 4090, winnow:e4b answers a five-question request in 89 ms end to end, and scores 0.722 on typed decisions against 0.738 for TypeSafe's hosted Jev. Smaller models such as laya answer in about 10 ms, and run well on a CPU.

89ms

winnow:e4b on Ollaya

RTX 4090, five questions · 0.722 accuracy

236–276ms

TypeSafe Jev

Hosted API, median request · 0.738 accuracy

Accuracy and speed of every model · typed-decisions accuracy, higher is better; latency, lower is better
ModelAccuracyLatency
  • TypeSafe Jevhosted API0.738236–276 ms
  • winnow:e4b0.72289 ms
  • kev:9b0.722498 ms
  • winnow:12b0.702131 ms
  • decider:4b0.680520 ms
  • kev:4b0.669354 ms
  • jevk5:4b0.625105 ms
  • decider:2b0.591190 ms
  • nli0.54820 ms
  • decider:0.8b0.506155 ms
  • gliclass0.47715 ms
  • kev:0.8b0.460128 ms
  • von0.44723 ms
  • laya:en0.36110 ms
  • clm:8bquestions cached0.357149 ms

Accuracy: the typed-decisions test split (400 states, 2,000 questions), argmax against the majority label, measured by Ollaya for each model; Jev's from Winnow's benchmark report on the same questions. laya:typed-decisions scores 0.766 but was fine-tuned on this dataset, so it is left out. clm caches questions and options: its latency is a new message whose questions are cached. Latency: median of a five-question request through the HTTP API on an NVIDIA RTX 4090; Jev: median request of the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), which includes the network. Setups differ, so read the latencies as an order-of-magnitude comparison.

Drop-in compatible

Speaks TypeSafe's API.

Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.

Request

# Point the TypeSafe SDK at Ollaya
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value works
export TYPESAFE_DEFAULT_MODEL=winnow:e4b

# …or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d '{
    "model": "winnow:e4b",
    "state": "Can I get an invoice for last month?",
    "questions": {
      "intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {
          "invoice": "Needs an invoice or receipt",
          "refund": "Wants money back",
          "other": "Anything else"
        }
      }
    }
  }'

Response

{
  "model": "winnow:e4b",
  "answers": {
    "intent": {
      "type": "choice",
      "choice": "invoice",
      "confidence": 0.9801,
      "probabilities": {
        "invoice": 0.9868,
        "refund": 0.0026,
        "other": 0.0106
      }
    }
  },
  "usage": {
    "input_tokens": 120,
    "output_tokens": 0
  }
}

TypeSafe compatibility guide

Open models

Open weights, ready to pull.

Pick by what you need: winnow:e4b balances accuracy and speed best, laya is the fastest and runs well on a CPU, kev and decider scale up to 9B and 4B, von reads up to 8,192 tokens, and qwen3guard screens text for safety. The models page shows each one’s accuracy and speed.

Your data stays yours

Private by default.

Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.

  • Local

    Runs on your machine with ONNX Runtime, on the CPU or an NVIDIA GPU. The server listens on 127.0.0.1 by default.

  • Open weights

    Weights come from their authors’ Hugging Face repositories, pinned to a commit and checked against sha256. Ollaya never re-hosts them, and the runtime is Apache-2.0.

  • No per-token fees

    Run as many decisions as your hardware can handle. No metering and no API bill.

  • Calibrated

    Probabilities you can put thresholds on. Each model ships its own calibration, and a Modelfile refits it on your labelled data.

Platforms

Runs where you work.

A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, Windows, WSL 2 or Docker takes a request down to milliseconds.

macOSApple silicon, macOS 14+Desktop appMenu bar app.dmgCommand lineInstall scriptGPUApple GPULaya and NLI on MLX
Windows10 and 11, x64Desktop appDesktop app.exe or .msiCommand linePowerShell scriptGPUNVIDIA, CUDA 13 or 12Command line
Linuxx86-64Desktop appDesktop appAppImage, .deb, .rpmCommand lineInstall scriptsystemd serviceGPUNVIDIA, CUDA 13 or 12
LinuxARM64Desktop appNot availableCommand lineInstall scriptsystemd serviceGPUCPU only
WSL 2Linux on WindowsDesktop appNot availableCommand lineInstall scriptSame as LinuxGPUNVIDIA, CUDA 13 or 12
Dockeramd64 and arm64Desktop appNot availableCommand lineImage on GHCRGPUNVIDIA, CUDA 13 or 12:cuda and :cuda12, amd64
Install for your platform

NVIDIA GPUs need driver R525 or newer; the install scripts fetch the CUDA libraries only when they find one. On a Mac, laya and nli run on the Apple GPU through MLX; other models, AMD and Intel GPUs, and the Windows and Linux desktop apps use the CPU.

Get up and running in minutes.

One binary, one command: ollaya run winnow:e4b.

macOS, Windows, Linux and Docker · Apache-2.0 · GitHub