Skip to content

snap

2 TagsUpdated Apache-2.0by logitlab

logitlab's snap1-2b: MiniCPM5-2B fine-tuned to read one option letter after the prompt of emnlmn's snap engine. Ollaya runs the author's Q8_0 GGUF on llama.cpp with snap's own prompt, on NVIDIA GPUs, Apple silicon or the CPU.

decisionfine-tunedgguf2b
ollaya run snap --preset triage "I was charged twice for my subscription this month and want a refund."

Models

View all
Name
snaplatest2.7 GB · 8192 ctx · English, it
snap:2b2.7 GB · 8192 ctx · English, it

Readme

snap1-2b by logitlab is MiniCPM5-2B fine-tuned for one job: a typed decision read from the option letters after one prompt. It is the model of snap, emnlmn's decision engine, and it is trained on snap's exact prompt. Ollaya reads the probability of each option letter as the next token. It never generates text.

Needs Ollaya 0.10.0 or newer, the first release that runs the snap-v1 prompt.

Models

TagWeightsTyped-decisions accuracyFive questions, RTX 4090
snap:latest, snap:2bsnap1-2b (MiniCPM5-2B), Q8_0 GGUF, 2.7 GB0.64868 ms

Typed-decisions accuracy is the argmax against the majority label on all 400 typed-decisions states, measured by Ollaya, with a calibration error (ECE) of 0.062 on snap's raw probabilities. The time is the median of 15 warm requests of the triage preset through the HTTP API. The author reports 0.655 with snap itself on the same file: the prompts are identical, and snap reads the answer letters slightly differently (below). The author also reports that the checkpoint never trained on typed-decisions.

Usage

ollaya run snap --preset triage "My order never arrived and support ignores me. Refund me today or I'm switching to your competitor."

Point any TypeSafe client at http://localhost:11435 and set the model to snap.

How it works

  • Prompt. Ollaya builds snap's prompt exactly as snap 0.5.0 does: identical on all 573 test prompts, token for token. A structured state is written in TOON, snap's compact rendering of JSON; the question goes before the state or after it as snap decides for the whole request.
  • Options. Each option is a letter, up to 26 per question, read in one pass. A yes/no question reads A) Yes, B) No.
  • Calibration. None: snap1-2b ships raw probabilities, and so does Ollaya.
  • Engine. llama.cpp v0.5.0, ggml-org's own build, on an NVIDIA GPU (CUDA), an Apple silicon GPU (Metal) or the CPU. About 3 GB of memory.
  • Parity. Ollaya's runner matches stock llama.cpp (llama-server of the same build, on the same file, with snap's own prompt tokens) on the CPU and on CUDA: the same decision on all 573 test questions, probabilities within 1.9e-6.

Limits

  • Options. Up to 26 per question. snap reads a wider choice as one yes/no question per option; Ollaya answers it with TOO_MANY_OPTIONS.
  • Readout. Ollaya reads each letter's own token in one pass per question; snap also pools the letter's space- and newline-prefixed tokens and reads the state once for all questions.
  • Context. 8,192 tokens per question; a longer prompt is rejected, not cut.
  • Languages. English and Italian.