Decision models explained: Jev, Kev, Tev1 and Solar Decide
A decision model turns a supplied state into a bounded answer: which team handles a request, whether evidence meets a policy, or which category fits a document. The application defines the possible answers before inference. This is useful where software needs a reliable answer shape instead of an open-ended reply.
Compare decision-model benchmark results →
Typed answers are not guaranteed correct answers
A classifier can select a valid option and still select the wrong one. Probabilities can help an application send uncertain cases for review, but calibration is a property measured across predictions. It does not certify one decision. Our ranking measures accuracy; it does not establish confidence calibration or unrestricted agent ability. TypeSafe documentation.
How the four implementations differ
Jev: TypeSafe's System One model returns typed choices, scores and yes/no probabilities. Its vendor describes parallel sampling and Reinforcement Learning for Calibrated Decisions. It gives up free text generation and reasoning explanations. Its exact parameter count and full architecture are not publicly disclosed. Jev introduction.
Kev 4B: Jared Palmer's public implementation adds a trained adapter and pointer head to Qwen3.5-4B-Base. The head selects supplied options rather than writing a chat reply. It is a separate implementation inspired by the decision interface. Kev source.
Tev1 4B Experimental: Together's supervised fine-tune retains Qwen's next-token output head. It uses a regular chat request to select one option letter, with thinking disabled. The published interface supports 2–24 choices. This differs from Jev's dedicated runtime; token log probabilities should not be described as calibrated confidence. Tev1 model card.
Solar Decide: Upstage exposes a System One decision endpoint on Solar Mini 4. It accepts state and typed questions, including choices, scores and yes/no answers, with a 512K context window. Solar Decide specification.
One small decision, explicit answer space
For a ticket saying an export fails, the application can supply a state and a Choice question with billing, technical, sales and security as options. The selected key determines the next application step. Keep transaction checks, counting and exact arithmetic in ordinary code; ask the model for the semantic judgment that code cannot express reliably.
{"state":"My CSV export fails.","questions":{"route":{"type":"choice","instructions":"Which team handles this request?","criteria":{"billing":"Payments","technical":"Product failures","sales":"Purchases","security":"Unauthorized access"}}}}How AI BENCHY evaluates decision models
We use OpenRouter's unified Decisions API. That wrapper also serves models with different underlying runtimes. Each of our four dedicated tests contains 12 independent cases: request routing, policy exceptions and missing evidence, support versus contradiction, and resistance to instructions embedded in untrusted records. The same state, questions and answer options go to every model. Hidden expected labels never enter its request.
Each model runs three repeats: 144 scored decisions. Accuracy determines the ranking; equal accuracy is ordered by measured suite cost. A perfect test attempt needs all 12 answers correct. Partial scores, execution errors and missing coverage remain distinct. Unfinished coverage receives no rank. The saved provider responses retain choices, probabilities when returned, usage and timing.
The dedicated suite does not change the general leaderboard. Older compatible benchmarks remain adapted selection tasks, with unsupported generation and tool-use tasks skipped. We skip answer spaces beyond Tev1's published 24-option limit and the 26-option limit reported by Solar Decide's OpenRouter route. Rejected requests remain in our private execution record; they do not count as wrong decisions. These four bounded tests describe this workload; they do not establish universal decision quality.