decision
2 TagsUpdated Apache-2.0by vLLM Semantic Router
Decision models by the vLLM Semantic Router contributors: a fully fine-tuned Qwen3.5 backbone plus an endpoint head that scores every option at its own last token against the question, in one forward pass per question. 16k-token rows.
decisionlong-context0.75b
Tags
| Name |
|---|
| decisionlatestSame as decision:eos.1.5 GB · 16384 ctx · English, Chinese |
| decision:eosDecision 1.0 Eos, fully fine-tuned Qwen3.5-0.8B: 17.49 on Decision Index 0.2.1.5 GB · 16384 ctx · English, Chinese |
Tags without a suffix load fp16 on a CUDA GPU and fp32 on CPU; add -fp16 or -fp32 to pin one precision.