Skip to content

snap:2b-q4_k_m

3 TagsUpdated 2B params8192 contextEnglish, itApache-2.0by logitlab

snap1-2b, Q4_K_M GGUF (snap's own default file, 1.6 GB): 0.660 on typed decisions; 63 ms for five questions on an RTX 4090.

2b
ollaya run snap:2b-q4_k_m --preset triage "I was charged twice for my subscription this month and want a refund."

Details

  • weights0bee5ed145f7 · 1.6 GBgguf · Q4_K - Medium · huggingface.co/logitlab/snap1-2b-GGUF/resolve/3932191…/snap1-2b-q4_k_m.gguf
  • decision33500e138657 · 1 KB{"engine": "llama", "family": "snap", "layout": "snap-v1", "gguf": {…}, …}
  • calibrationd25f1f5a12b7 · 173 B{"temperature": [1.0, 1.0, 1.0]}
  • license50f3b1f52af3 · 10 KBsnap1-2b by logitlab (https://huggingface.co/logitlab/snap1-2b-GGUF), openbmb/MiniCPM5-2B (Apache-2.0) fine-tuned with a LoRA and merged, Apache-2.0. Its prompt is snap's (https://github.com/emnlmn/snap, MIT), ported to Ollaya's runtime.

Every layer is checked against its sha256 when it is pulled. Weights and tokenizers download from the model author's Hugging Face repository at a pinned commit; Ollaya never re-hosts them.

snap readme and all models