Results
What we measure on our own GPUs and CPUs: how accurate each model is, how well its confidence is calibrated, how fast it runs, and whether it gives the same answers as its authors' code. All numbers, with the raw data.
Newest measurement: 2026-10-01
73,720
benchmark answers scored against human labels
41,352
questions checked against the authors' own code
25
models measured
1,989
requests timed on our GPUs and CPUs
Test bench
The machines behind these numbers.
Two desktops with NVIDIA GPUs and a Mac mini. Every number on this page was measured on one of them, with the same model files that ollaya pull fetches from each author's repository.
RTX 5090
RTX 5090 desktop
- GPU
- RTX 5090, 32 GB
- CPU
- AMD Ryzen Threadripper 3960X (24 cores, 48 threads)
- Memory
- 64 GB
- System
- Windows 11 Enterprise LTSC, Ubuntu 24.04 under WSL2
- Driver
- NVIDIA 596.21 (CUDA 13.2)
Measured here
- Bespoke Labs' public benchmark, with Ollama on the same GPU
- Speed of every model on the RTX 5090 and on the Threadripper
- CUDA on Blackwell (sm_120)
RTX 4090
RTX 4090 desktop
- GPU
- RTX 4090, 24 GB
- CPU
- Intel Core i9-13900K (24 cores, 32 threads)
- Memory
- 64 GB
- System
- Windows 11 Pro, Ubuntu 24.04 under WSL2 (24 of the 32 threads)
- Driver
- NVIDIA 616.92
Measured here
- Parity on CUDA, under Linux and Windows, and on Vulkan
- Speed of every model on the RTX 4090 and on the i9-13900K
- Typed-decisions quality of every model
- Windows: CUDA, Vulkan and the desktop app
M4 Pro
Mac mini (M4 Pro)
- GPU
- Apple M4 Pro GPU, 16 cores
- CPU
- Apple M4 Pro (12 cores: 8 performance, 4 efficiency)
- Memory
- 24 GB
- System
- macOS 27
Measured here
- Parity on the Apple GPU (MLX) and the Apple CPU
- The macOS app and the menu bar
Software: Ollaya 0.8.0, which runs ONNX models on ONNX Runtime (with NVIDIA's CUDA 13 libraries on the GPUs), GGUF models on llama.cpp build b11146, and laya and nli on the Apple GPU through MLX. Community benches, measured by contributors with the same tools: Intel Arc 140T laptop (Intel Core Ultra 9 285H), by MauricioPerera; Radeon RX 9070 desktop (AMD Ryzen 9 9950X), by solarpush.
Accuracy and speed
Side by side with Ollama, on the same GPU.
Ollaya is inspired by Ollama, which now serves two decision models of its own. Both ran Bespoke Labs' public decision benchmark: 3,880 human-labeled questions from 13 datasets, scored with Bespoke's own code, one request at a time on one RTX 5090.
- Ollaya 0.8.0
- Ollama 0.35.0
Calibration
Confidence you can put a threshold on.
Expected calibration error (ECE) measures how far a model's confidence is from how often it is right. On the same Nimble weights, Ollaya's ECE is 0.022 and Ollama's 0.122: Ollaya applies each model's fitted temperature, Ollama returns the raw softmax.
- Ollaya 0.8.0
- Ollama 0.35.0
Accuracy by dataset
Where each model is strong, dataset by dataset.
Fact checking, intent, reading comprehension, paraphrase, inference, toxicity, safety, helpfulness, summary quality and biomedical questions. Each cell is the share of questions answered like the human label, in percent.
| VitaminC | MASSIVE en | MASSIVE de | BoolQ | SQuAD 2 | PAWS | MultiNLI | Civil Comments | Aegis 2 | HelpSteer 2 | SummEval rel. | SummEval cons. | PubMedQA | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| winnow:12b | 75 | 87 | 85 | 87 | 88 | 83 | 82 | 87 | 82 | 41 | 50 | 88 | 71 |
| decider:4b | 74 | 86 | 84 | 87 | 79 | 77 | 93 | 87 | 80 | 45 | 38 | 82 | 71 |
| kev:9b | 73 | 85 | 84 | 85 | 74 | 80 | 91 | 84 | 80 | 39 | 46 | 82 | 76 |
| nimble (Ollama) | 79 | 84 | 83 | 86 | 74 | 74 | 90 | 78 | 82 | 35 | 48 | 83 | 77 |
| nimble:9b | 77 | 84 | 81 | 85 | 82 | 80 | 83 | 84 | 81 | 34 | 40 | 85 | 76 |
| tev1:4b (Ollama) | 75 | 86 | 83 | 85 | 77 | 82 | 92 | 74 | 77 | 37 | 49 | 80 | 75 |
| kev:4b | 72 | 86 | 86 | 85 | 78 | 82 | 92 | 82 | 81 | 34 | 37 | 81 | 72 |
| winnow:e4b | 70 | 85 | 81 | 82 | 82 | 81 | 83 | 75 | 80 | 38 | 45 | 85 | 69 |
| decider:2b | 64 | 82 | 79 | 80 | 78 | 72 | 86 | 84 | 72 | 43 | 35 | 67 | 73 |
| decider:0.8b | 68 | 82 | 78 | 79 | 72 | 59 | 82 | 87 | 62 | 40 | 14 | 84 | 62 |
| tev1:0.8b (Ollama) | 69 | 79 | 69 | 78 | 71 | 64 | 76 | 77 | 57 | 33 | 14 | 84 | 60 |
| jevk5:4b | 68 | 0 | 0 | 85 | 80 | 81 | 84 | 83 | 80 | 42 | 44 | 85 | 74 |
| kev:0.8b | 66 | 75 | 66 | 79 | 66 | 66 | 80 | 84 | 69 | 29 | 14 | 55 | 45 |
| decision:eos | 62 | 82 | 75 | 68 | 87 | 67 | 77 | 65 | 69 | 26 | 22 | 29 | 37 |
| laya:multilingual | 76 | 62 | 44 | 75 | 60 | 77 | 85 | 92 | 60 | 26 | 13 | 30 | 52 |
| laya:en | 79 | 69 | 41 | 84 | 62 | 58 | 88 | 94 | 43 | 16 | 6 | 4 | 50 |
| von:1.1 | 76 | 69 | 27 | 72 | 50 | 58 | 33 | 76 | 63 | 14 | 32 | 8 | 54 |
| nli:deberta-v3-large | 26 | 60 | 55 | 75 | 47 | 60 | 31 | 89 | 61 | 17 | 23 | 3 | 50 |
| gliclass:large | 12 | 57 | 49 | 58 | 51 | 54 | 32 | 80 | 57 | 27 | 46 | 8 | 32 |
| Model | Accuracy | Choice | Yes/no | Score | ECE | Median | Rejected |
|---|---|---|---|---|---|---|---|
| winnow:12b | 0.773 | 0.800 | 0.854 | 0.592 | 0.141 | 60 ms | 0 |
| decider:4b | 0.756 | 0.817 | 0.820 | 0.549 | 0.043 | 188 ms | 0 |
| kev:9b | 0.753 | 0.816 | 0.807 | 0.556 | 0.058 | 229 ms | 0 |
| nimble (Ollama) | 0.749 | 0.825 | 0.790 | 0.553 | 0.122 | 210 ms | 0 |
| nimble:9b | 0.748 | 0.800 | 0.826 | 0.529 | 0.022 | 310 ms | 0 |
| tev1:4b (Ollama) | 0.747 | 0.820 | 0.792 | 0.552 | 0.075 | 406 ms | 0 |
| kev:4b | 0.745 | 0.815 | 0.816 | 0.508 | 0.062 | 196 ms | 0 |
| winnow:e4b | 0.734 | 0.776 | 0.798 | 0.558 | 0.071 | 43 ms | 0 |
| decider:2b | 0.703 | 0.767 | 0.773 | 0.481 | 0.087 | 98 ms | 0 |
| decider:0.8b | 0.669 | 0.745 | 0.719 | 0.459 | 0.061 | 88 ms | 0 |
| tev1:0.8b (Ollama) | 0.639 | 0.707 | 0.693 | 0.436 | 0.129 | 61 ms | 0 |
| jevk5:4b | 0.621 | 0.454 | 0.818 | 0.570 | 0.040 | 26 ms | 700 |
| kev:0.8b | 0.611 | 0.664 | 0.729 | 0.325 | 0.035 | 106 ms | 0 |
| decision:eos | 0.591 | 0.667 | 0.714 | 0.257 | 0.087 | 96 ms | 0 |
| laya:multilingual | 0.579 | 0.639 | 0.729 | 0.228 | 0.155 | 14 ms | 27 |
| laya:en | 0.532 | 0.652 | 0.679 | 0.087 | 0.106 | 17 ms | 453 |
| von:1.1 | 0.485 | 0.516 | 0.636 | 0.181 | 0.098 | 16 ms | 0 |
| nli:deberta-v3-large | 0.458 | 0.443 | 0.664 | 0.141 | 0.196 | 14 ms | 326 |
| gliclass:large | 0.434 | 0.366 | 0.600 | 0.269 | 0.284 | 14 ms | 28 |
Speed on our machines
Every model, on every machine we have.
The triage preset (five questions) on a short customer message, through the HTTP API, one request at a time: what a client sees. A different message for each request, so nothing is answered from a cache.
- RTX 5090
- RTX 4090
- Threadripper 3960X CPU
- i9-13900K CPU
Windows: CUDA and Vulkan
Vulkan brings GGUF models to any GPU.
Vulkan runs on NVIDIA, AMD and Intel GPUs without the 1.4 GB CUDA libraries (pull request #27). On the same RTX 4090 under Windows 11, Ollaya's runtime matched llama.cpp's own server on every question with both backends; CUDA is faster.
- CUDA
- Vulkan
Parity
The same answers as the authors' own code.
Before a model ships, Ollaya's runtime is compared question by question with a reference on each device: the authors' code in fp32 for ONNX models, and llama.cpp's own server of the same build for GGUF models. Every decision must be the same and every option score within 0.001.
| Model | CUDA | Vulkan | Apple GPU (MLX) | CPU |
|---|---|---|---|---|
| clm:8b | ✓ passed 480 q | · | · | ✓ passed 480 q |
| cygnet:12b | ✓ passed 502 q | ✓ passed 502 q | · | · |
| decider:2b | ✓ passed 479 q | · | · | · |
| decider:2b-vision | ✓ passed 404 q | · | · | ✓ passed 404 q |
| decider:4b | ✓ passed 479 q | · | · | ✓ passed 479 q |
| decision:eos | ✓ passed 466 q | · | · | ✓ passed 466 q |
| jeb:27b | ✓ passed 494 q | · | · | · |
| jeb:4b | ✓ passed 494 q | · | · | · |
| jeb:9b | ✓ passed 494 q | ✓ passed 494 q | · | · |
| jeeves:9b | ✓ passed 430 q | · | · | · |
| jevk5:4b | ✓ passed 593 q | · | · | · |
| kev:0.8b | ✓ passed 480 q | · | · | ✓ passed 480 q |
| kev:4b | ✓ passed 480 q | · | · | ✓ passed 480 q |
| kev:9b | ✓ passed 480 q | · | · | ✓ passed 480 q |
| laya:en | ✓ passed 483 q | · | ✓ passed 2,383 q | ✓ passed 2,383 q |
| laya:multilingual | ✓ passed 2,383 q | · | ✓ passed 2,383 q | · |
| nimble:9b | ✓ passed 492 q | · | · | · |
| nli:deberta-v3-large | ✓ passed | · | · | · |
| nli:modernbert-large | · | · | ✓ passed 483 q | ✓ passed 483 q |
| qwen3guard:0.6b | ✓ passed 296 q | · | · | ✓ passed 296 q |
| von:1.1 | ✓ passed 485 q | · | ✗ outside 485 q | ✓ passed 485 q |
| winnow:12b | ✓ passed 505 q | · | · | · |
| winnow:e4b | ✓ passed 505 q | ✓ passed 505 q | · | ✓ passed 505 q |
The count is the number of questions compared. "Outside": measured, and over the gate on at least one question; that device does not run the model (details in the raw data). Per family: docs/families.
- CPU
- Vulkan
Raw data
Every number, as the tools wrote it.
Each chart on this page is drawn from these files. The tools that wrote them are in the repository (convert/ollaya_convert/bench_public.py, bench_latency.py and the parity examples), so you can run the same measurements on your own hardware.
- 2026-09-30-public-benchmark-rtx5090-cuda.json
- 2026-10-01-latency-rtx4090-cpu-extra.json
- 2026-10-01-latency-rtx4090-cpu.json
- 2026-10-01-latency-rtx4090-cuda-extra.json
- 2026-10-01-latency-rtx4090-cuda.json
- 2026-10-01-latency-rtx5090-cpu.json
- 2026-10-01-latency-rtx5090-cuda.json
- 2026-10-01-parity-rtx4090-windows.json
- parity-docs.json