Decision model ranking
Four tests. 48 decisions. Three repeats per model.
How decision models work (English) →| # | Model | Accuracy | Attempts | Perfect attempts | Time / request | Suite cost | Execution errors | Routing | Policy | Evidence | Injection |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Jev 1.13 | 100.00% | 12/12 | 12/12 | 0.38s | $0.00108 | 0 | 100.0% | 100.0% | 100.0% | 100.0% |
| 2 | Kev 4B | 89.58% | 12/12 | 0/12 | 0.98s | $0.00066 | 0 | 91.7% | 91.7% | 91.7% | 83.3% |
| 3 | Solar Decide | 89.58% | 12/12 | 3/12 | 10.09s | $0.00612 | 0 | 100.0% | 91.7% | 91.7% | 75.0% |
| 4 | Tev1 4B Experimental | 87.50% | 12/12 | 3/12 | 0.88s | $0.00370 | 0 | 91.7% | 100.0% | 91.7% | 66.7% |
Ranked by correct decisions; cost breaks ties. A perfect attempt needs all 12 answers correct. Missing coverage receives no rank.
Cost covers all 12 requests. Time measures a whole request, including all its questions. Providers can process these questions differently. These results are separate from the general leaderboard.