Skip to content

decision

2 TagsUpdated Apache-2.0by vLLM Semantic Router

Decision models by the vLLM Semantic Router contributors: a fully fine-tuned Qwen3.5 backbone plus an endpoint head that scores every option at its own last token against the question, in one forward pass per question. 16k-token rows.

decisionlong-context0.75b

Tags

Name
decisionlatestSame as decision:eos.1.5 GB · 16384 ctx · English, Chinese
decision:eosDecision 1.0 Eos, fully fine-tuned Qwen3.5-0.8B: 17.49 on Decision Index 0.2.1.5 GB · 16384 ctx · English, Chinese

Tags without a suffix load fp16 on a CUDA GPU and fp32 on CPU; add -fp16 or -fp32 to pin one precision.