Skip to content

Results

What we measure on our own GPUs and CPUs: how accurate each model is, how well its confidence is calibrated, how fast it runs, and whether it gives the same answers as its authors' code. All numbers, with the raw data.

Newest measurement: 2026-10-01

73,720

benchmark answers scored against human labels

41,352

questions checked against the authors' own code

25

models measured

1,989

requests timed on our GPUs and CPUs

Test bench

The machines behind these numbers.

Two desktops with NVIDIA GPUs and a Mac mini. Every number on this page was measured on one of them, with the same model files that ollaya pull fetches from each author's repository.

RTX 5090

RTX 5090 desktop

GPU
RTX 5090, 32 GB
CPU
AMD Ryzen Threadripper 3960X (24 cores, 48 threads)
Memory
64 GB
System
Windows 11 Enterprise LTSC, Ubuntu 24.04 under WSL2
Driver
NVIDIA 596.21 (CUDA 13.2)

Measured here

  • Bespoke Labs' public benchmark, with Ollama on the same GPU
  • Speed of every model on the RTX 5090 and on the Threadripper
  • CUDA on Blackwell (sm_120)

RTX 4090

RTX 4090 desktop

GPU
RTX 4090, 24 GB
CPU
Intel Core i9-13900K (24 cores, 32 threads)
Memory
64 GB
System
Windows 11 Pro, Ubuntu 24.04 under WSL2 (24 of the 32 threads)
Driver
NVIDIA 616.92

Measured here

  • Parity on CUDA, under Linux and Windows, and on Vulkan
  • Speed of every model on the RTX 4090 and on the i9-13900K
  • Typed-decisions quality of every model
  • Windows: CUDA, Vulkan and the desktop app

M4 Pro

Mac mini (M4 Pro)

GPU
Apple M4 Pro GPU, 16 cores
CPU
Apple M4 Pro (12 cores: 8 performance, 4 efficiency)
Memory
24 GB
System
macOS 27

Measured here

  • Parity on the Apple GPU (MLX) and the Apple CPU
  • The macOS app and the menu bar

Software: Ollaya 0.8.0, which runs ONNX models on ONNX Runtime (with NVIDIA's CUDA 13 libraries on the GPUs), GGUF models on llama.cpp build b11146, and laya and nli on the Apple GPU through MLX. Community benches, measured by contributors with the same tools: Intel Arc 140T laptop (Intel Core Ultra 9 285H), by MauricioPerera; Radeon RX 9070 desktop (AMD Ryzen 9 9950X), by solarpush.

Accuracy and speed

Side by side with Ollama, on the same GPU.

Ollaya is inspired by Ollama, which now serves two decision models of its own. Both ran Bespoke Labs' public decision benchmark: 3,880 human-labeled questions from 13 datasets, scored with Bespoke's own code, one request at a time on one RTX 5090.

Accuracy against latency · up and to the left is better
  • Ollaya 0.8.0
  • Ollama 0.35.0
Accuracy against median latency on Bespoke Labs' public benchmark, RTX 5090, Ollaya and Ollama0.40.50.60.70.8102050100200500median latency per question (ms), log scaleaccuracykev:9bnimblenimble:9bdecider:4bkev:4btev1:4bdecider:0.8btev1:0.8bkev:0.8bwinnow:12bwinnow:e4bdecider:2bdecision:eosvon:1.1gliclass:largelaya:multilingual
Accuracy: the mean over the 13 datasets. Latency: the median request (one question), HTTP included. The most accurate model, winnow:12b, scores 0.773 at 60 ms; Ollama's best, nimble, 0.749 at 210 ms. Models that reject more than 5 % of the questions (option limits, long states) are left out of the chart; the table under the heatmap lists them.

Calibration

Confidence you can put a threshold on.

Expected calibration error (ECE) measures how far a model's confidence is from how often it is right. On the same Nimble weights, Ollaya's ECE is 0.022 and Ollama's 0.122: Ollaya applies each model's fitted temperature, Ollama returns the raw softmax.

Calibration error, all 3,880 questions · ten-bin ECE of the top probability; lower is better
  • Ollaya 0.8.0
  • Ollama 0.35.0
nimble:9b0.022
kev:0.8b0.035
jevk5:4b18 % rejected0.040
decider:4b0.043
kev:9b0.058
decider:0.8b0.061
kev:4b0.062
winnow:e4b0.071
tev1:4b (Ollama)0.075
decision:eos0.087
decider:2b0.087
von:1.10.098
laya:en12 % rejected0.106
nimble (Ollama)0.122
tev1:0.8b (Ollama)0.129
winnow:12b0.141
laya:multilingual0.155
nli:deberta-v3-large8 % rejected0.196
gliclass:large0.284

Accuracy by dataset

Where each model is strong, dataset by dataset.

Fact checking, intent, reading comprehension, paraphrase, inference, toxicity, safety, helpfulness, summary quality and biomedical questions. Each cell is the share of questions answered like the human label, in percent.

Accuracy per dataset (%) · darker is higher; rows sorted by the mean; Ollama's rows in orange
VitaminCMASSIVE enMASSIVE deBoolQSQuAD 2PAWSMultiNLICivil CommentsAegis 2HelpSteer 2SummEval rel.SummEval cons.PubMedQA
winnow:12b75878587888382878241508871
decider:4b74868487797793878045388271
kev:9b73858485748091848039468276
nimble (Ollama)79848386747490788235488377
nimble:9b77848185828083848134408576
tev1:4b (Ollama)75868385778292747737498075
kev:4b72868685788292828134378172
winnow:e4b70858182828183758038458569
decider:2b64827980787286847243356773
decider:0.8b68827879725982876240148462
tev1:0.8b (Ollama)69796978716476775733148460
jevk5:4b680085808184838042448574
kev:0.8b66756679666680846929145545
decision:eos62827568876777656926222937
laya:multilingual76624475607785926026133052
laya:en796941846258889443166450
von:1.17669277250583376631432854
nli:deberta-v3-large2660557547603189611723350
gliclass:large1257495851543280572746832
Accuracy by question type, calibration error, latency and rejected questions per model
ModelAccuracyChoiceYes/noScoreECEMedianRejected
winnow:12b0.7730.8000.8540.5920.14160 ms0
decider:4b0.7560.8170.8200.5490.043188 ms0
kev:9b0.7530.8160.8070.5560.058229 ms0
nimble (Ollama)0.7490.8250.7900.5530.122210 ms0
nimble:9b0.7480.8000.8260.5290.022310 ms0
tev1:4b (Ollama)0.7470.8200.7920.5520.075406 ms0
kev:4b0.7450.8150.8160.5080.062196 ms0
winnow:e4b0.7340.7760.7980.5580.07143 ms0
decider:2b0.7030.7670.7730.4810.08798 ms0
decider:0.8b0.6690.7450.7190.4590.06188 ms0
tev1:0.8b (Ollama)0.6390.7070.6930.4360.12961 ms0
jevk5:4b0.6210.4540.8180.5700.04026 ms700
kev:0.8b0.6110.6640.7290.3250.035106 ms0
decision:eos0.5910.6670.7140.2570.08796 ms0
laya:multilingual0.5790.6390.7290.2280.15514 ms27
laya:en0.5320.6520.6790.0870.10617 ms453
von:1.10.4850.5160.6360.1810.09816 ms0
nli:deberta-v3-large0.4580.4430.6640.1410.19614 ms326
gliclass:large0.4340.3660.6000.2690.28414 ms28

Speed on our machines

Every model, on every machine we have.

The triage preset (five questions) on a short customer message, through the HTTP API, one request at a time: what a client sees. A different message for each request, so nothing is answered from a cache.

Five questions, median request · further left is faster; GPUs in blue, CPUs in orange, one mark per machine
  • RTX 5090
  • RTX 4090
  • Threadripper 3960X CPU
  • i9-13900K CPU
Median latency of five-question requests per model, on each GPU and CPU we measured101001k10k100kmilliseconds for five questions, log scalelaya:multilinguallaya:multilingual on RTX 5090: 14 mslaya:multilingual on RTX 4090: 7 mslaya:multilingual on Threadripper 3960X CPU: 472 mslaya:multilingual on i9-13900K CPU: 308 mslaya:enlaya:en on RTX 5090: 16 mslaya:en on RTX 4090: 11 mslaya:en on Threadripper 3960X CPU: 915 mslaya:en on i9-13900K CPU: 619 msgliclass:largegliclass:large on RTX 5090: 16 msgliclass:large on RTX 4090: 17 msgliclass:large on Threadripper 3960X CPU: 1872 msgliclass:large on i9-13900K CPU: 563 msnli:modernbert-largenli:modernbert-large on RTX 5090: 17 msnli:modernbert-large on RTX 4090: 23 msnli:modernbert-large on Threadripper 3960X CPU: 1848 msnli:modernbert-large on i9-13900K CPU: 595 msnli:deberta-v3-largenli:deberta-v3-large on RTX 5090: 18 msnli:deberta-v3-large on RTX 4090: 22 msnli:deberta-v3-large on Threadripper 3960X CPU: 2405 msnli:deberta-v3-large on i9-13900K CPU: 781 msvon:1.1von:1.1 on RTX 5090: 21 msvon:1.1 on RTX 4090: 23 msvon:1.1 on Threadripper 3960X CPU: 2445 msvon:1.1 on i9-13900K CPU: 753 msqwen3guard:0.6bqwen3guard:0.6b on RTX 5090: 35 msqwen3guard:0.6b on RTX 4090: 34 msqwen3guard:0.6b on Threadripper 3960X CPU: 1442 msqwen3guard:0.6b on i9-13900K CPU: 802 msjeb:4bjeb:4b on RTX 5090: 98 msjeb:4b on RTX 4090: 95 msjeb:4b on Threadripper 3960X CPU: 5296 msjeb:4b on i9-13900K CPU: 7485 msclm:8bclm:8b on RTX 5090: 100 msclm:8b on RTX 4090: 137 msclm:8b on Threadripper 3960X CPU: 4972 msclm:8b on i9-13900K CPU: 3952 msjevk5:4bjevk5:4b on RTX 5090: 106 msjevk5:4b on RTX 4090: 103 msjevk5:4b on Threadripper 3960X CPU: 7109 msjevk5:4b on i9-13900K CPU: 9466 mswinnow:e4bwinnow:e4b on RTX 5090: 121 mswinnow:e4b on RTX 4090: 103 mswinnow:e4b on Threadripper 3960X CPU: 3993 mswinnow:e4b on i9-13900K CPU: 6119 msjeb:9bjeb:9b on RTX 5090: 113 msjeb:9b on RTX 4090: 124 msjeb:9b on Threadripper 3960X CPU: 9506 msjeb:9b on i9-13900K CPU: 11894 mskev:0.8bkev:0.8b on RTX 5090: 145 mskev:0.8b on RTX 4090: 133 mskev:0.8b on Threadripper 3960X CPU: 2836 mskev:0.8b on i9-13900K CPU: 1267 msdecider:2b-visiondecider:2b-vision on RTX 5090: 151 msdecider:2b-vision on RTX 4090: 142 msdecider:2b-vision on Threadripper 3960X CPU: 3123 msdecider:2b-vision on i9-13900K CPU: 2015 mswinnow:12bwinnow:12b on RTX 5090: 157 mswinnow:12b on RTX 4090: 158 mswinnow:12b on Threadripper 3960X CPU: 10928 mswinnow:12b on i9-13900K CPU: 14824 mscygnet:12bcygnet:12b on RTX 5090: 169 mscygnet:12b on RTX 4090: 197 mscygnet:12b on Threadripper 3960X CPU: 20636 mscygnet:12b on i9-13900K CPU: 23583 msdecider:0.8bdecider:0.8b on RTX 5090: 177 msdecider:0.8b on RTX 4090: 169 msdecider:0.8b on Threadripper 3960X CPU: 3035 msdecider:0.8b on i9-13900K CPU: 1811 msdecider:2bdecider:2b on RTX 5090: 208 msdecider:2b on RTX 4090: 216 msdecider:2b on Threadripper 3960X CPU: 4397 msdecider:2b on i9-13900K CPU: 3357 msdecision:eosdecision:eos on RTX 5090: 215 msdecision:eos on RTX 4090: 210 msdecision:eos on Threadripper 3960X CPU: 3402 msdecision:eos on i9-13900K CPU: 2022 msjeb:27bjeb:27b on RTX 5090: 261 msjeb:27b on RTX 4090: 301 msjeb:27b on Threadripper 3960X CPU: 19985 msjeb:27b on i9-13900K CPU: 25215 mskev:4bkev:4b on RTX 5090: 327 mskev:4b on RTX 4090: 360 mskev:4b on Threadripper 3960X CPU: 8298 mskev:4b on i9-13900K CPU: 6081 mskev:9bkev:9b on RTX 5090: 375 mskev:9b on RTX 4090: 506 mskev:9b on Threadripper 3960X CPU: 12117 mskev:9b on i9-13900K CPU: 10475 msdecider:4bdecider:4b on RTX 5090: 440 msdecider:4b on RTX 4090: 514 msdecider:4b on Threadripper 3960X CPU: 10391 msdecider:4b on i9-13900K CPU: 8779 msjeeves:9bjeeves:9b on RTX 5090: 668 msjeeves:9b on RTX 4090: 864 msjeeves:9b on Threadripper 3960X CPU: 20785 msjeeves:9b on i9-13900K CPU: 19521 msnimble:9bnimble:9b on RTX 5090: 1789 msnimble:9b on RTX 4090: 2276 msnimble:9b on Threadripper 3960X CPU: 59779 msnimble:9b on i9-13900K CPU: 50730 ms
Ollaya 0.8.0, Linux under WSL2. Each model: one load, 5 untimed requests, then 20 timed ones (on the CPU 2 and 10). A missing dot: the model did not run there or is still being measured.

Windows: CUDA and Vulkan

Vulkan brings GGUF models to any GPU.

Vulkan runs on NVIDIA, AMD and Intel GPUs without the 1.4 GB CUDA libraries (pull request #27). On the same RTX 4090 under Windows 11, Ollaya's runtime matched llama.cpp's own server on every question with both backends; CUDA is faster.

Five questions, median, RTX 4090 under Windows · the runtime without HTTP; lower is better
  • CUDA
  • Vulkan
winnow:e4bCUDA103 ms
winnow:e4bVulkan149 ms
jeb:9bCUDA128 ms
jeb:9bVulkan250 ms
cygnet:12bCUDA213 ms
cygnet:12bVulkan862 ms

Parity

The same answers as the authors' own code.

Before a model ships, Ollaya's runtime is compared question by question with a reference on each device: the authors' code in fp32 for ONNX models, and llama.cpp's own server of the same build for GGUF models. Every decision must be the same and every option score within 0.001.

Runtime parity per model and device: passed or outside the gate, with the number of questions compared
ModelCUDAVulkanApple GPU (MLX)CPU
clm:8b✓ 480 q··✓ 480 q
cygnet:12b✓ 502 q✓ 502 q··
decider:2b✓ 479 q···
decider:2b-vision✓ 404 q··✓ 404 q
decider:4b✓ 479 q··✓ 479 q
decision:eos✓ 466 q··✓ 466 q
jeb:27b✓ 494 q···
jeb:4b✓ 494 q···
jeb:9b✓ 494 q✓ 494 q··
jeeves:9b✓ 430 q···
jevk5:4b✓ 593 q···
kev:0.8b✓ 480 q··✓ 480 q
kev:4b✓ 480 q··✓ 480 q
kev:9b✓ 480 q··✓ 480 q
laya:en✓ 483 q·✓ 2,383 q✓ 2,383 q
laya:multilingual✓ 2,383 q·✓ 2,383 q·
nimble:9b✓ 492 q···
nli:deberta-v3-large✓···
nli:modernbert-large··✓ 483 q✓ 483 q
qwen3guard:0.6b✓ 296 q··✓ 296 q
von:1.1✓ 485 q·✗ 485 q✓ 485 q
winnow:12b✓ 505 q···
winnow:e4b✓ 505 q✓ 505 q·✓ 505 q

The count is the number of questions compared. "Outside": measured, and over the gate on at least one question; that device does not run the model (details in the raw data). Per family: docs/families.

Same prompt, different backends · largest difference in an option's log-probability against the RTX 4090's CUDA numbers
  • CPU
  • Vulkan
winnow:e4bcpu (x86-64), 501 of 505 decisions the same0.28
winnow:e4bvulkan (RTX 4090), 501 of 505 decisions the same0.32
jeb:9bvulkan (RTX 4090), 492 of 494 decisions the same0.33
cygnet:12bcpu (i9-13900K, Windows), 496 of 502 decisions the same2.76
cygnet:12bvulkan (RTX 4090), 497 of 502 decisions the same1.40
The same reference prompts on different devices, compared with the RTX 4090's CUDA goldens: how much backends differ. Measured, not gated.

Raw data

Every number, as the tools wrote it.

Each chart on this page is drawn from these files. The tools that wrote them are in the repository (convert/ollaya_convert/bench_public.py, bench_latency.py and the parity examples), so you can run the same measurements on your own hardware.