jeb
4 TagsUpdated Apache-2.0by frontier-infra
Decision models by Jason Brashear and AINode: LoRAs merged into Qwen3.5 and Qwen3.8, published as GGUF. Jeb reads the option labels' logits after AINode's own decision prompt, with a fitted temperature per question type; Ollaya runs the authors' files on llama.cpp.
ollaya run jeb --preset triage "I was charged twice for my subscription this month and want a refund."curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "jeb",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}'# Already using a TypeSafe SDK? Set TYPESAFE_BASE_URL=http://localhost:11435 instead.
import requests
response = requests.post(
"http://localhost:11435/api/decide",
json={
"model": "jeb",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
},
)
answers = response.json()["answers"]
print(answers["department"]["choice"], answers["refund"]["noul"])const response = await fetch("http://localhost:11435/api/decide", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "jeb",
state: "I was charged twice for my subscription this month and want a refund.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoices and refunds",
technical: "Bugs, errors and outages",
account: "Login, profile and settings"
}
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?"
}
}
}),
});
const { answers } = await response.json();
console.log(answers.department.choice, answers.refund.noul);Models
View all| Name |
|---|
| jeblatest9.8 GB · 4096 ctx · English |
| jeb:27b16.8 GB · 4096 ctx · English |
| jeb:4b4.6 GB · 4096 ctx · English |
| jeb:9b9.8 GB · 4096 ctx · English |
Each model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.
Readme
Jebadiah (Jeb) is a family of decision models by Jason Brashear, built with AINode and released under Apache-2.0: rank-16 LoRAs merged into Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B. The authors publish them as GGUF files, and Ollaya runs those files as they are, on llama.cpp. Jeb reads a question as AINode's decision prompt (the state, the question and lettered options) and returns the probability of each option label as the next token. It never generates text.
Needs Ollaya 0.8.0 or newer, the first release that runs the
jebadiah-v1prompt.
Models
| Tag | Base | Weights | Five questions, RTX 4090 |
|---|---|---|---|
jeb:4b | Qwen3.5-4B, v2 | Q8_0 GGUF, 4.5 GB | 96 ms |
jeb:latest, jeb:9b | Qwen3.5-9B, v2 | Q8_0 GGUF, 9.8 GB | 124 ms |
jeb:27b | Qwen3.8-27B | Q4_K_M GGUF, 17 GB | 309 ms |
The 4B and 9B files are Q8_0, the ones the authors checked against the bf16 weights (256 and 257 of 260 held-out answers the same). The 27B is Q4_K_M so that it fits a 24 GB GPU. Latency is the triage preset (five questions) on a short message through the HTTP API, at the median of 15 warm requests.
On typed-decisions Jeb scores about 0.79 to 0.80, but its training data includes the typed-decisions train split, so that number is not comparable with the zero-shot models and the model list leaves it out. The authors publish held-out results in getainode/jebadiah.
Usage
ollaya run jeb --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."Point any TypeSafe client at http://localhost:11435 and set the model to jeb, jeb:4b or jeb:27b.
How it works
- Prompt. Ollaya builds the prompt exactly as the authors' renderer does (
scripts/jebadiah_prompt.py, which is AINode's own/v1/systemonerenderer): identical on all 1,349 test prompts. - Options. A yes/no question reads
trueasAandfalseasB; a choice reads its options in order; a score its levels. Labels runA..Z,AA,AB, ..., up to 255 options. - Calibration. The authors' fitted temperature per question type (
temperatures.json). - One pass per question. Qwen3.5's recurrent layers cannot share a cached state prefix, so each question is evaluated from the start. The same request always returns the same probabilities.
- Engine. llama.cpp v0.5.0, ggml-org's own build, on an NVIDIA GPU (CUDA), an Apple silicon GPU (Metal) or the CPU.
- Parity. Ollaya's runner matches stock llama.cpp (
llama-serverof the same build, on the same files) on CUDA: the same decision on all 494 test questions for each model, probabilities within 2.4e-6.
Limits
- Prompt. Up to 2,048 tokens per question, the state included, the budget the authors serve with. A longer prompt is rejected; the authors' renderer cuts the state instead.
- Instructions. Must be a non-empty string.
- Control tokens. Text you send can never become one of Qwen's control tokens.