Skip to content

nimble

2 TagsUpdated Apache-2.0by Bespoke Labs

Decision model by Bespoke Labs: a LoRA on Qwen3.5-9B trained on contrastive pairs. Nimble reads the whole request as a JSON schema and scores each option by the next-token logit of its code; Ollaya applies the author's temperature, so its probabilities are calibrated. Up to 255 options.

decisionfine-tuned9b
ollaya run nimble --preset triage "I was charged twice for my subscription this month and want a refund."

Models

View all
Name
nimblelatest19.5 GB · 8192 ctx · English
nimble:9b19.5 GB · 8192 ctx · English

Each model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.

Readme

Nimble is a decision model by Bespoke Labs: a LoRA adapter on Qwen3.5-9B, trained on contrastive pairs (two examples that differ by one fact that flips the answer), released under Apache-2.0. For each question it reads the whole request as a JSON schema and scores every option by the next-token logit of its code (A, B, ...). It never generates text.

Needs Ollaya 0.8.0 or newer, the first release that runs the nimble-codes-v1 layout.

Models

TagBaseCheckpointParamsTyped-decisions accuracy
nimble:latest, nimble:9bQwen3.5-9BBespoke-Nimble-9B-v2 (2026-09-23)9B0.665

Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya. On Bespoke Labs' public benchmark (13 human-labeled datasets, 3,880 questions), see the comparison with Ollama.

Usage

ollaya run nimble --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."

Point any TypeSafe client at http://localhost:11435 and set the model to nimble.

Speed

  • RTX 4090, through the HTTP API: the triage preset (five questions) on a short message takes 2,297 ms at the median of 15 warm requests (about 3,000 input tokens: each of the five rows carries the whole request).
  • State length: every question's row repeats the whole request (the state and every question's schema), so the cost grows with the number of questions times the request length.
  • CPU: a 9B model in fp32 arithmetic belongs on a GPU.

How it works

  • Prompt. Ollaya builds Nimble's prompt exactly as the author's code does (serving_schema.prepare_prompts in the model repository), and maps TypeSafe questions to Nimble's schema the way the author's own /v1/systemone server does: noul becomes a true/false field, choice an enum of its labels, score an enum of its levels.
  • Calibrated. Probabilities use the author's recommended temperature, 2.179 (temperature_config.json). A plain softmax of the code logits, what Ollama returns, is overconfident.
  • Up to 255 options. Questions with more than 26 options use the author's serving extension: two-letter codes and the "short code" prompt.
  • Weights. Qwen's Qwen3.5-9B shards and Bespoke Labs' adapter download from Hugging Face, pinned to a commit and verified by sha256. The adapter is applied at run time, not merged, so every file stays byte for byte the authors'. Ollaya hosts only the ONNX graph (14 MB).
  • Parity. Ollaya's Rust runtime matches the author's code in fp32 on CUDA: identical token rows, the same rejections and the same decision on all 492 test questions, probabilities within 6.5e-6.

Limits

  • Context. Each question's row, the whole request included, can be up to 8,192 tokens; a longer one is rejected, never cut.
  • State. The state must not be empty; instructions and choice labels must not be blank.
  • Memory. The weights stay BF16 in memory: about 18 GB, so a 24 GB GPU.