snap
2 TagsUpdated Apache-2.0by logitlab
logitlab's snap1-2b: MiniCPM5-2B fine-tuned to read one option letter after the prompt of emnlmn's snap engine. Ollaya runs the author's Q8_0 GGUF on llama.cpp with snap's own prompt, on NVIDIA GPUs, Apple silicon or the CPU.
ollaya run snap --preset triage "I was charged twice for my subscription this month and want a refund."curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "snap",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}'# Already using a TypeSafe SDK? Set TYPESAFE_BASE_URL=http://localhost:11435 instead.
import requests
response = requests.post(
"http://localhost:11435/api/decide",
json={
"model": "snap",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
},
)
answers = response.json()["answers"]
print(answers["department"]["choice"], answers["refund"]["noul"])const response = await fetch("http://localhost:11435/api/decide", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "snap",
state: "I was charged twice for my subscription this month and want a refund.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoices and refunds",
technical: "Bugs, errors and outages",
account: "Login, profile and settings"
}
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?"
}
}
}),
});
const { answers } = await response.json();
console.log(answers.department.choice, answers.refund.noul);Models
View allReadme
snap1-2b by logitlab is MiniCPM5-2B fine-tuned for one job: a typed decision read from the option letters after one prompt. It is the model of snap, emnlmn's decision engine, and it is trained on snap's exact prompt. Ollaya reads the probability of each option letter as the next token. It never generates text.
Needs Ollaya 0.10.0 or newer, the first release that runs the
snap-v1prompt.
Models
| Tag | Weights | Typed-decisions accuracy | Five questions, RTX 4090 |
|---|---|---|---|
snap:latest, snap:2b | snap1-2b (MiniCPM5-2B), Q8_0 GGUF, 2.7 GB | 0.648 | 68 ms |
Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya, with a calibration error (ECE) of 0.062 on snap's raw probabilities. The time is the median of 15 warm requests of the triage preset through the HTTP API. The author reports 0.655 with snap itself on the same file: the prompts are identical, and snap reads the answer letters slightly differently (below). The author also reports that the checkpoint never trained on typed-decisions.
Usage
ollaya run snap --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."Point any TypeSafe client at http://localhost:11435 and set the model to snap.
How it works
- Prompt. Ollaya builds snap's prompt exactly as snap 0.5.0 does: identical on all 573 test prompts, token for token. A structured state is written in TOON, snap's compact rendering of JSON; the question goes before the state or after it as snap decides for the whole request.
- Options. Each option is a letter, up to 26 per question, read in one pass. A yes/no question reads
A) Yes,B) No. - Calibration. None: snap1-2b ships raw probabilities, and so does Ollaya.
- Engine. llama.cpp v0.5.0, ggml-org's own build, on an NVIDIA GPU (CUDA), an Apple silicon GPU (Metal) or the CPU. About 3 GB of memory.
- Parity. Ollaya's runner matches stock llama.cpp (
llama-serverof the same build, on the same file, with snap's own prompt tokens) on the CPU and on CUDA: the same decision on all 573 test questions, probabilities within 1.9e-6.
Limits
- Options. Up to 26 per question. snap reads a wider choice as one yes/no question per option; Ollaya answers it with
TOO_MANY_OPTIONS. - Readout. Ollaya reads each letter's own token in one pass per question; snap also pools the letter's space- and newline-prefixed tokens and reads the state once for all questions.
- Context. 8,192 tokens per question; a longer prompt is rejected, not cut.
- Languages. English and Italian.