The average score is effectively tied at 3.0 vs 3.0. Command A+ (high) has the lower benchmark cost at $0.000 vs $0.002. Kev 4B is faster at 5.63s vs 86.95s, with pass rates of 0.0% vs 16.7%.
Last updated at:
2026-09-28
Compared models
Rank
#362
Total Output Tokens
0
Response Time (avg)
86.95s
Total Cost
$0.000
Rank
#366
Total Output Tokens
28,061
Response Time (avg)
5.63s
Total Cost
$0.002
Recommended modelKev 4B
It has the best score here (3.0), while responding about 15.4x faster than Command A+ (high).
Kev 4BKev 4Bnone36/66 attempts. Choices supplied by adapters. Supports 12/22 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-28
Kev 4BKev 4Bnone36/66 attempts. Choices supplied by adapters. Supports 12/22 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-28
Score
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#362
#366
Reliability
0.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
5.1Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
66/66
36/66
Tests Correct
A test is fully passed only if every run passed for that test.API error: 22Response Time (avg)86.95sResponse Time (max)114.87sResponse Time (total)1913.01sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 9Response Time (avg)5.63sResponse Time (max)17.79sResponse Time (total)67.61sA test is fully passed only if every run passed for that test.…
Attempt pass rate
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
16.7%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
66Total Runs…
36Total Runs…
Cost per result
0.000
0.060
Total Cost
$0.000
$0.002
Input Price
$0.300 / 1MInput Price…
$0.042 / 1MInput Price…
Output Price
$1.500 / 1MOutput Price…
$0.000 / 1MOutput Price…
Total Input Tokens
0Total Input Tokens…
42,804Total Input Tokens…
Output Tokens
0Output Tokens…
28,061Output Tokens…
Reasoning Tokens
0Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
86.95sResponse Time (avg)…
5.63sResponse Time (avg)…
Response Time (max)
114.87sResponse Time (max)…
17.79sResponse Time (max)…
Response Time (total)
1913.01sResponse Time (total)…
67.61sResponse Time (total)…
Parameters
218B total (25B active)
4B
Availability
Open source
Open source
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#362 Command A+
high
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok
#366 Jared Palmer: Kev 4B
none
No showcase result has been generated for this model yet.
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 4Response Time (avg)85.94sResponse Time (max)91.77sResponse Time (total)343.74sA test is fully passed only if every run passed for that test.…
2.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)3.65sResponse Time (max)9.32sResponse Time (total)10.95sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)92.74sResponse Time (max)101.39sResponse Time (total)278.22sA test is fully passed only if every run passed for that test.…
4.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)771msResponse Time (max)954msResponse Time (total)1.54sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)80.65sResponse Time (max)82.20sResponse Time (total)161.31sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)86.20sResponse Time (max)89.21sResponse Time (total)172.39sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)9.09sResponse Time (max)17.37sResponse Time (total)18.18sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)67.89sResponse Time (max)70.95sResponse Time (total)203.68sA test is fully passed only if every run passed for that test.…
5.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
33.3%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)6.34sResponse Time (max)17.79sResponse Time (total)19.02sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)99.30sResponse Time (max)99.30sResponse Time (total)99.30sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)110.33sResponse Time (max)114.87sResponse Time (total)220.65sA test is fully passed only if every run passed for that test.…
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)17.27sResponse Time (max)17.27sResponse Time (total)17.27sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)86.20sResponse Time (max)92.42sResponse Time (total)258.61sA test is fully passed only if every run passed for that test.…
2.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)646msResponse Time (max)646msResponse Time (total)646msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)84.61sResponse Time (max)84.61sResponse Time (total)84.61sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)90.50sResponse Time (max)90.50sResponse Time (total)90.50sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…