jeeves
2 TagsUpdated Apache-2.0by PostHog
Decision model by PostHog: Qwen3.5-9B with a LoRA merged in and a pointer head that scores every option at its own marker. Ollaya runs it without its reasoning chain, in one forward pass per question, calibrated with the authors' temperature.
ollaya run jeeves --preset triage "I was charged twice for my subscription this month and want a refund."curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "jeeves",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}'# Already using a TypeSafe SDK? Set TYPESAFE_BASE_URL=http://localhost:11435 instead.
import requests
response = requests.post(
"http://localhost:11435/api/decide",
json={
"model": "jeeves",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
},
)
answers = response.json()["answers"]
print(answers["department"]["choice"], answers["refund"]["noul"])const response = await fetch("http://localhost:11435/api/decide", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "jeeves",
state: "I was charged twice for my subscription this month and want a refund.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoices and refunds",
technical: "Bugs, errors and outages",
account: "Login, profile and settings"
}
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?"
}
}
}),
});
const { answers } = await response.json();
console.log(answers.department.choice, answers.refund.noul);Models
View allEach model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.
Readme
Jeeves is a decision model by PostHog: Qwen3.5-9B with a LoRA merged in and a pointer head that scores every option at its own marker, released under Apache-2.0. Upstream, Jeeves can write a reasoning chain before it answers; Ollaya runs it without the chain, in one forward pass per question. It never generates text.
Needs Ollaya 0.8.0 or newer, the first release that runs the
jeeves-markers-v1layout.
Models
| Tag | Base | Params | Typed-decisions accuracy | Five questions, RTX 4090 |
|---|---|---|---|---|
jeeves:latest, jeeves:9b | Qwen3.5-9B, LoRA merged | 9B | 0.680 | 838 ms |
Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya, with a calibration error (ECE) of 0.031. The authors report 0.804 on their own test split without thinking (0.840 with it), with a calibration error of 0.021 at the fitted temperature.
Usage
ollaya run jeeves --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."Point any TypeSafe client at http://localhost:11435 and set the model to jeeves.
How it works
- Rows. Ollaya builds each question exactly as the authors' encoder does: the state, the question and the options between marker tokens, an empty thought, then the options again and a decide marker. The pointer head scores each option at its marker against the decide marker.
- No thinking. Jeeves' reasoning mode generates a chain first, which Ollaya's one-pass runtime does not do. Without it the authors report 0.804 against 0.840 on their test split, at a fraction of the latency.
- Calibrated. One fitted temperature, 1.859, from the model's
export.json. - Weights. PostHog's five BF16 shards and
head.ptdownload from Hugging Face, pinned to a commit and verified by sha256. Ollaya hosts only the ONNX graph (13 MB). - Parity. Ollaya's Rust runtime matches the authors' own code in fp32 on CUDA: identical token rows and the same decision on all 430 test questions, probabilities within 1.4e-5.
Limits
- Questions. As upstream, a question takes only
type,instructionsandcriteria, and choice criteria must be an object of options. - Rows. Each question's row, the state included, can be up to 8,192 tokens.
- Memory. About 18 GB with the weights kept BF16: a 24 GB GPU.