Skip to content

jeeves

2 TagsUpdated Apache-2.0by PostHog

Decision model by PostHog: Qwen3.5-9B with a LoRA merged in and a pointer head that scores every option at its own marker. Ollaya runs it without its reasoning chain, in one forward pass per question, calibrated with the authors' temperature.

decisionfine-tuned9b
ollaya run jeeves --preset triage "I was charged twice for my subscription this month and want a refund."

Models

View all
Name
jeeveslatest17.9 GB · 8192 ctx · English
jeeves:9b17.9 GB · 8192 ctx · English

Each model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.

Readme

Jeeves is a decision model by PostHog: Qwen3.5-9B with a LoRA merged in and a pointer head that scores every option at its own marker, released under Apache-2.0. Upstream, Jeeves can write a reasoning chain before it answers; Ollaya runs it without the chain, in one forward pass per question. It never generates text.

Needs Ollaya 0.8.0 or newer, the first release that runs the jeeves-markers-v1 layout.

Models

TagBaseParamsTyped-decisions accuracyFive questions, RTX 4090
jeeves:latest, jeeves:9bQwen3.5-9B, LoRA merged9B0.680838 ms

Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya, with a calibration error (ECE) of 0.031. The authors report 0.804 on their own test split without thinking (0.840 with it), with a calibration error of 0.021 at the fitted temperature.

Usage

ollaya run jeeves --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."

Point any TypeSafe client at http://localhost:11435 and set the model to jeeves.

How it works

  • Rows. Ollaya builds each question exactly as the authors' encoder does: the state, the question and the options between marker tokens, an empty thought, then the options again and a decide marker. The pointer head scores each option at its marker against the decide marker.
  • No thinking. Jeeves' reasoning mode generates a chain first, which Ollaya's one-pass runtime does not do. Without it the authors report 0.804 against 0.840 on their test split, at a fraction of the latency.
  • Calibrated. One fitted temperature, 1.859, from the model's export.json.
  • Weights. PostHog's five BF16 shards and head.pt download from Hugging Face, pinned to a commit and verified by sha256. Ollaya hosts only the ONNX graph (13 MB).
  • Parity. Ollaya's Rust runtime matches the authors' own code in fp32 on CUDA: identical token rows and the same decision on all 430 test questions, probabilities within 1.4e-5.

Limits

  • Questions. As upstream, a question takes only type, instructions and criteria, and choice criteria must be an object of options.
  • Rows. Each question's row, the state included, can be up to 8,192 tokens.
  • Memory. About 18 GB with the weights kept BF16: a 24 GB GPU.