Skip to content

jeb

4 TagsUpdated Apache-2.0by frontier-infra

Decision models by Jason Brashear and AINode: LoRAs merged into Qwen3.5 and Qwen3.8, published as GGUF. Jeb reads the option labels' logits after AINode's own decision prompt, with a fitted temperature per question type; Ollaya runs the authors' files on llama.cpp.

decisionfine-tunedgguf4b9b27b
ollaya run jeb --preset triage "I was charged twice for my subscription this month and want a refund."

Models

View all
Name
jeblatest9.8 GB · 4096 ctx · English
jeb:27b16.8 GB · 4096 ctx · English
jeb:4b4.6 GB · 4096 ctx · English
jeb:9b9.8 GB · 4096 ctx · English

Each model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.

Readme

Jebadiah (Jeb) is a family of decision models by Jason Brashear, built with AINode and released under Apache-2.0: rank-16 LoRAs merged into Qwen3.5-4B, Qwen3.5-9B and Qwen3.8-27B. The authors publish them as GGUF files, and Ollaya runs those files as they are, on llama.cpp. Jeb reads a question as AINode's decision prompt (the state, the question and lettered options) and returns the probability of each option label as the next token. It never generates text.

Needs Ollaya 0.8.0 or newer, the first release that runs the jebadiah-v1 prompt.

Models

TagBaseWeightsFive questions, RTX 4090
jeb:4bQwen3.5-4B, v2Q8_0 GGUF, 4.5 GB96 ms
jeb:latest, jeb:9bQwen3.5-9B, v2Q8_0 GGUF, 9.8 GB124 ms
jeb:27bQwen3.8-27BQ4_K_M GGUF, 17 GB309 ms

The 4B and 9B files are Q8_0, the ones the authors checked against the bf16 weights (256 and 257 of 260 held-out answers the same). The 27B is Q4_K_M so that it fits a 24 GB GPU. Latency is the triage preset (five questions) on a short message through the HTTP API, at the median of 15 warm requests.

On typed-decisions Jeb scores about 0.79 to 0.80, but its training data includes the typed-decisions train split, so that number is not comparable with the zero-shot models and the model list leaves it out. The authors publish held-out results in getainode/jebadiah.

Usage

ollaya run jeb --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."

Point any TypeSafe client at http://localhost:11435 and set the model to jeb, jeb:4b or jeb:27b.

How it works

  • Prompt. Ollaya builds the prompt exactly as the authors' renderer does (scripts/jebadiah_prompt.py, which is AINode's own /v1/systemone renderer): identical on all 1,349 test prompts.
  • Options. A yes/no question reads true as A and false as B; a choice reads its options in order; a score its levels. Labels run A..Z, AA, AB, ..., up to 255 options.
  • Calibration. The authors' fitted temperature per question type (temperatures.json).
  • One pass per question. Qwen3.5's recurrent layers cannot share a cached state prefix, so each question is evaluated from the start. The same request always returns the same probabilities.
  • Engine. llama.cpp v0.5.0, ggml-org's own build, on an NVIDIA GPU (CUDA), an Apple silicon GPU (Metal) or the CPU.
  • Parity. Ollaya's runner matches stock llama.cpp (llama-server of the same build, on the same files) on CUDA: the same decision on all 494 test questions for each model, probabilities within 2.4e-6.

Limits

  • Prompt. Up to 2,048 tokens per question, the state included, the budget the authors serve with. A longer prompt is rejected; the authors' renderer cuts the state instead.
  • Instructions. Must be a non-empty string.
  • Control tokens. Text you send can never become one of Qwen's control tokens.