decima:small
4 TagsUpdated 122M params512 context100+ languagesApache-2.0by A. M. Madani
Decima-small 1.1 (multilingual-e5-small, 122M), fp32: 0.432 on typed decisions; 146 ms for five questions on a CPU, the fastest there.
multilingual122m
ollaya run decima:small --preset triage "I was charged twice for my subscription this month and want a refund."curl http://localhost:11435/api/decide \
-H "Content-Type: application/json" \
-d '{
"model": "decima:small",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
}'# Already using a TypeSafe SDK? Set TYPESAFE_BASE_URL=http://localhost:11435 instead.
import requests
response = requests.post(
"http://localhost:11435/api/decide",
json={
"model": "decima:small",
"state": "I was charged twice for my subscription this month and want a refund.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoices and refunds",
"technical": "Bugs, errors and outages",
"account": "Login, profile and settings"
}
},
"refund": {
"type": "noul",
"instructions": "Is the customer asking for a refund?"
}
}
},
)
answers = response.json()["answers"]
print(answers["department"]["choice"], answers["refund"]["noul"])const response = await fetch("http://localhost:11435/api/decide", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "decima:small",
state: "I was charged twice for my subscription this month and want a refund.",
questions: {
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: {
billing: "Payments, invoices and refunds",
technical: "Bugs, errors and outages",
account: "Login, profile and settings"
}
},
refund: {
type: "noul",
instructions: "Is the customer asking for a refund?"
}
}
}),
});
const { answers } = await response.json();
console.log(answers.department.choice, answers.refund.noul);Details
- graphf83c59768e53 · 3 MB
onnx · decima · 122M · fp32 - weights2bc43026b30d · 471 MB
huggingface.co/amyrmahdy/decima-small/resolve/2e7f4d0…/pytorch/encoder/model.safetensors - weights5f2a149e4a4f · 19 MB
huggingface.co/amyrmahdy/decima-small/resolve/2e7f4d0…/pytorch/head.safetensors - tokenizer255f5e32cb32 · 17 MB
huggingface.co/amyrmahdy/decima-small/resolve/2e7f4d0…/pytorch/encoder/tokenizer.json - decisionf9f995481ae1 · 3 KB
{"engine": "onnx", "family": "decima", "encoder": "", "layout": "decima-late-interaction-v1", …} - calibrationeb3ca5d61970 · 291 B
{"temperature": [0.936, 1.0, 0.936]} - licensed2a0668879d4 · 11 KB
Decima-small, Decima-base and Decima-agent by A. M. Madani (https://huggingface.co/amyrmahdy, https://github.com/amyrmahdy/decima), Apache-2.0.
Every layer is checked against its sha256 when it is pulled. Weights and tokenizers download from the model author's Hugging Face repository at a pinned commit; Ollaya never re-hosts them.