cygnet
2 TagsUpdated Apache-2.0by blockbrain-ai
Frozen Gemma 4 12B IT with blockbrain-ai's Cygnet prompt: the options as letters, one answer slot, and one calibration temperature. No fine-tuning; Ollaya runs ggml-org's Q8_0 GGUF on llama.cpp.
ollaya run cygnet --preset triage "I was charged twice for my subscription this month and want a refund."curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "cygnet",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}'# Already using a TypeSafe SDK? Set TYPESAFE_BASE_URL=http://localhost:11435 instead.
import requests
response = requests.post(
"http://localhost:11435/api/decide",
json={
"model": "cygnet",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
},
)
answers = response.json()["answers"]
print(answers["department"]["choice"], answers["refund"]["noul"])const response = await fetch("http://localhost:11435/api/decide", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "cygnet",
state: "I was charged twice for my subscription this month and want a refund.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoices and refunds",
technical: "Bugs, errors and outages",
account: "Login, profile and settings"
}
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?"
}
}
}),
});
const { answers } = await response.json();
console.log(answers.department.choice, answers.refund.noul);Models
View all| Name |
|---|
| cygnetlatest12.7 GB · 16384 ctx · 100+ languages |
| cygnet:12b12.7 GB · 16384 ctx · 100+ languages |
Each model carries fp16 and fp32 graphs over one weights file, and loads fp16 on a CUDA GPU and fp32 on CPU.
Readme
Cygnet by blockbrain-ai is not a fine-tune: it is Google's Gemma 4 12B IT, unchanged, plus a system message, the options as letters, one answer slot and one calibration temperature. Ollaya reads the probability of each option letter as the next token. It never generates text.
Needs Ollaya 0.8.0 or newer, the first release that runs the
cygnet-v1prompt.
Models
| Tag | Weights | Typed-decisions accuracy | Five questions, RTX 4090 |
|---|---|---|---|
cygnet:latest, cygnet:12b | Gemma 4 12B IT, Q8_0 GGUF, 12.7 GB | 0.683 | 202 ms |
Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya, with a calibration error (ECE) of 0.148 at Cygnet's temperature (0.289 without it). The authors report 203 of 231 on JevBench's public items with BF16 weights on vLLM; Ollaya runs ggml-org's Q8_0 conversion of the same Gemma revision on llama.cpp, which is close but not bit for bit the same model.
Usage
ollaya run cygnet --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."Point any TypeSafe client at http://localhost:11435 and set the model to cygnet.
How it works
- Prompt. Ollaya builds Cygnet's prompt exactly as its decision server does (
shim/decision_server.pyandshim/cygnet_shim.py): identical on all 1,364 test prompts. A structured state is shown as indented JSON, the rendering Cygnet was measured with. - Options. A yes/no question reads
falseasAandtrueasB, the order Cygnet was measured in. Up to 20 options per question, read in one pass. - Calibration. Temperature 3.4 over the letters, Cygnet's own.
- Engine. llama.cpp v0.5.0, ggml-org's own build, on an NVIDIA GPU (CUDA), an Apple silicon GPU (Metal) or the CPU. A 16 GB GPU holds it.
- Parity. Ollaya's runner matches stock llama.cpp (
llama-serverof the same build, on the same file) on CUDA: the same decision on all 502 test questions, probabilities within 4.1e-7.
Limits
- Options. Up to 20 per question. Cygnet's server reads wider questions in groups; Ollaya answers them with
TOO_MANY_OPTIONS. - Context. 16,384 tokens per question; a longer prompt is rejected, not cut.