Navigate
Advertise here

Qwen3.5-122B-A10B vs Grok 4.3 (medium)

The average score is effectively tied at 6.4 vs 6.5. Qwen3.5-122B-A10B has the lower benchmark cost at $0.308 vs $0.989. Qwen3.5-122B-A10B is faster at 14.79s vs 44.53s, with pass rates of 36.2% vs 63.8%.

Last updated at: 2026-10-01

Compared models

Rank
#201
Total Output Tokens
97,112
Response Time (avg)
14.79s
Total Cost
$0.308
Rank
#199
Total Output Tokens
232,542
Response Time (avg)
44.53s
Total Cost
$0.989
Recommended model Qwen3.5-122B-A10B

It has the best score here (6.4), while costing about 3.2x less than Grok 4.3 (medium).

Detailed comparison

Metric Qwen3.5-122B-A10B Qwen3.5-122B-A10B none Release: 2026-02-24 Grok 4.3 Grok 4.3 medium Release: 2026-05-01
Score 6.4 6.5
Rank #201 #199
Reliability 10.0 10.0
Consistency 9.3 8.3
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 36.2% 63.8%
Flaky tests 2 5
Total Runs 69 69
Cost per result 4.452 8.235
Total Cost $0.308 $0.989
Input Price $0.260 / 1M $1.250 / 1M
Output Price $2.080 / 1M $2.500 / 1M
Total Input Tokens 405,921 325,471
Output Tokens 97,112 14,867
Reasoning Tokens 0 217,675
Response Time (avg) 14.79s 44.53s
Response Time (max) 212.63s 216.69s
Response Time (total) 340.25s 1024.25s
Parameters 122B total (10B active) ~1.5T total (~150B active)
Availability Open source Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#201 Qwen3.5-122B-A10B

none
Cost
$0.016
Time
44.5s
Tokens
6,431 tok

#199 SpaceXAI: Grok 4.3

medium
Cost
$0.009
Time
19.0s
Tokens
3,661 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.5-122B-A10B 3.7 7.0 22.2% 1 2.77s 7,913 693 0
Grok 4.3 5.9 7.7 44.4% 1 41.23s 8,340 1,028 31,226

Quick Compare

Switch Comparison Pair