Navigate
Advertise here

Command A+ vs Qwen3.5-9B (medium)

Qwen3.5-9B (medium) leads on average score with 3.8 vs 3.0. Command A+ has the lower benchmark cost at $0.000 vs $0.062. Command A+ is faster at 81.36s vs 107.51s, with pass rates of 0.0% vs 25.8%.

Last updated at: 2026-09-23

Compared models

Rank
#356
Total Output Tokens
0
Response Time (avg)
81.36s
Total Cost
$0.000
Rank
#346
Total Output Tokens
413,442
Response Time (avg)
107.51s
Total Cost
$0.062
Recommended model Qwen3.5-9B (medium)

It has the strongest score in this comparison (3.8) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric Command A+ Command A+ none Release: 2026-09-23 Qwen3.5-9B Qwen3.5-9B medium Release: 2026-03-02
Score 3.0 3.8
Rank #356 #346
Reliability 0.0 6.0
Consistency 10.0 8.1
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 0.0% 25.8%
Flaky tests 0 5
Total Runs 66 66
Cost per result 0.000 2.061
Total Cost $0.000 $0.062
Input Price $0.300 / 1M $0.100 / 1M
Output Price $1.500 / 1M $0.150 / 1M
Total Input Tokens 0 17,195
Output Tokens 0 46,735
Reasoning Tokens 0 366,707
Response Time (avg) 81.36s 107.51s
Response Time (max) 106.37s 569.97s
Response Time (total) 1789.98s 1720.15s
Parameters 218B total (25B active) 9B
Availability Open source Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#356 Command A+

none
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok

#346 Qwen3.5-9B

medium
Cost
$0.001
Time
35.9s
Tokens
3,030 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Command A+ 3.0 10.0 0.0% 0 80.82s 0 0 0
Qwen3.5-9B 2.9 10.0 0.0% 0 100.88s 2,396 7,890 41,129

Quick Compare

Switch Comparison Pair