Command A+ (high) leads on average score with 3.2 vs 1.9. D1 has the lower benchmark cost at $0.003 vs $0.079. D1 is faster at 508ms vs 89.96s, with pass rates of 0.0% vs 11.6%.
Last updated at:
2026-10-02
Compared models
Rank
#372
Total Output Tokens
28,663
Response Time (avg)
89.96s
Total Cost
$0.079
Rank
#382
Total Output Tokens
0
Response Time (avg)
508ms
Total Cost
$0.003
Recommended modelCommand A+ (high)
It has the strongest score in this comparison (3.2) and the best overall balance of cost and response time across all 2 models.
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
Score
3.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#372
#382
Reliability
3.5First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
4.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
69/69
36/69
Tests Correct
A test is fully passed only if every run passed for that test.API error: 22Wrong answer: 1Response Time (avg)89.96sResponse Time (max)156.00sResponse Time (total)2069.01sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 10Response Time (avg)508msResponse Time (max)820msResponse Time (total)6.10sA test is fully passed only if every run passed for that test.…
Attempt pass rate
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
11.6%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
69Total Runs…
36Total Runs…
Cost per result
0.000
0.112
Total Cost
$0.079
$0.003
Input Price
$0.300 / 1MInput Price…
$0.040 / 1MInput Price…
Output Price
$1.500 / 1MOutput Price…
$0.000 / 1MOutput Price…
Cache Read Price
$0.150 / 1MCache Read Price…
$0.040 / 1MCache Read Price…
Cache Write Price
N/ACache Write Price…
N/ACache Write Price…
Total Input Tokens
117,966Total Input Tokens…
55,956Total Input Tokens…
Output Tokens
5,149Output Tokens…
0Output Tokens…
Reasoning Tokens
23,514Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
89.96sResponse Time (avg)…
508msResponse Time (avg)…
Response Time (max)
156.00sResponse Time (max)…
820msResponse Time (max)…
Response Time (total)
2069.01sResponse Time (total)…
6.10sResponse Time (total)…
Parameters
218B total (25B active)
-
Availability
Open source
Closed
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#372 Command A+
high
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok
#382 LiquidAI: D1
none
No showcase result has been generated for this model yet.
3.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
9.6Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)156.00sResponse Time (max)156.00sResponse Time (total)156.00sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 4Response Time (avg)85.94sResponse Time (max)91.77sResponse Time (total)343.74sA test is fully passed only if every run passed for that test.…
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)431msResponse Time (max)505msResponse Time (total)1.29sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)92.74sResponse Time (max)101.39sResponse Time (total)278.22sA test is fully passed only if every run passed for that test.…
2.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)456msResponse Time (max)577msResponse Time (total)911msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)80.65sResponse Time (max)82.20sResponse Time (total)161.31sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)86.20sResponse Time (max)89.21sResponse Time (total)172.39sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)418msResponse Time (max)518msResponse Time (total)836msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)67.89sResponse Time (max)70.95sResponse Time (total)203.68sA test is fully passed only if every run passed for that test.…
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)555msResponse Time (max)589msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)99.30sResponse Time (max)99.30sResponse Time (total)99.30sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)110.33sResponse Time (max)114.87sResponse Time (total)220.65sA test is fully passed only if every run passed for that test.…
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)820msResponse Time (max)820msResponse Time (total)820msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)86.20sResponse Time (max)92.42sResponse Time (total)258.61sA test is fully passed only if every run passed for that test.…
1.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)570msResponse Time (max)570msResponse Time (total)570msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)84.61sResponse Time (max)84.61sResponse Time (total)84.61sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)90.50sResponse Time (max)90.50sResponse Time (total)90.50sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…