Command A+ (medium) leads on average score with 3.4 vs 1.9. D1 has the lower benchmark cost at $0.003 vs $0.064. D1 is faster at 508ms vs 91.72s, with pass rates of 0.0% vs 11.6%.
Last updated at:
2026-10-02
Compared models
Rank
#367
Total Output Tokens
18,795
Response Time (avg)
91.72s
Total Cost
$0.064
Rank
#382
Total Output Tokens
0
Response Time (avg)
508ms
Total Cost
$0.003
Recommended modelCommand A+ (medium)
It has the strongest score in this comparison (3.4) and the best overall balance of cost and response time across all 2 models.
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
D1D1none36/69 attempts. Choices supplied by adapters. Supports 12/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-10-02
Score
3.4Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.9Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#367
#382
Reliability
3.5First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
4.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
69/69
36/69
Tests Correct
A test is fully passed only if every run passed for that test.API error: 22Wrong answer: 1Response Time (avg)91.72sResponse Time (max)112.76sResponse Time (total)2109.54sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 10Response Time (avg)508msResponse Time (max)820msResponse Time (total)6.10sA test is fully passed only if every run passed for that test.…
Attempt pass rate
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
11.6%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
69Total Runs…
36Total Runs…
Cost per result
0.000
0.112
Total Cost
$0.064
$0.003
Input Price
$0.300 / 1MInput Price…
$0.040 / 1MInput Price…
Output Price
$1.500 / 1MOutput Price…
$0.000 / 1MOutput Price…
Cache Read Price
$0.150 / 1MCache Read Price…
$0.040 / 1MCache Read Price…
Cache Write Price
N/ACache Write Price…
N/ACache Write Price…
Total Input Tokens
118,771Total Input Tokens…
55,956Total Input Tokens…
Output Tokens
4,127Output Tokens…
0Output Tokens…
Reasoning Tokens
14,668Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
91.72sResponse Time (avg)…
508msResponse Time (avg)…
Response Time (max)
112.76sResponse Time (max)…
820msResponse Time (max)…
Response Time (total)
2109.54sResponse Time (total)…
6.10sResponse Time (total)…
Parameters
218B total (25B active)
-
Availability
Open source
Closed
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#367 Command A+
medium
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok
#382 LiquidAI: D1
none
No showcase result has been generated for this model yet.
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)112.76sResponse Time (max)112.76sResponse Time (total)112.76sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 4Response Time (avg)94.10sResponse Time (max)102.33sResponse Time (total)376.41sA test is fully passed only if every run passed for that test.…
4.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)431msResponse Time (max)505msResponse Time (total)1.29sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)93.13sResponse Time (max)94.47sResponse Time (total)279.38sA test is fully passed only if every run passed for that test.…
2.2Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)456msResponse Time (max)577msResponse Time (total)911msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)89.35sResponse Time (max)95.40sResponse Time (total)178.70sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)92.87sResponse Time (max)97.01sResponse Time (total)185.73sA test is fully passed only if every run passed for that test.…
6.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)418msResponse Time (max)518msResponse Time (total)836msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)77.87sResponse Time (max)91.53sResponse Time (total)233.61sA test is fully passed only if every run passed for that test.…
3.6Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
7.2Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
22.2%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 3Response Time (avg)555msResponse Time (max)589msResponse Time (total)1.66sA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)80.14sResponse Time (max)80.14sResponse Time (total)80.14sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 2Response Time (avg)103.03sResponse Time (max)105.69sResponse Time (total)206.06sA test is fully passed only if every run passed for that test.…
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)820msResponse Time (max)820msResponse Time (total)820msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 3Response Time (avg)92.07sResponse Time (max)99.99sResponse Time (total)276.20sA test is fully passed only if every run passed for that test.…
1.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)570msResponse Time (max)570msResponse Time (total)570msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)78.42sResponse Time (max)78.42sResponse Time (total)78.42sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.API error: 1Response Time (avg)102.12sResponse Time (max)102.12sResponse Time (total)102.12sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…