Navigate
Advertise here

GPT-6 Sol vs Qwen3.5-122B-A10B (medium)

The average score is effectively tied at 7.2 vs 7.1. GPT-6 Sol has the lower benchmark cost at $0.280 vs $1.047. GPT-6 Sol is faster at 6.84s vs 65.67s, with pass rates of 66.7% vs 71.2%.

Last updated at: 2026-09-23

Compared models

Rank
#148
Total Output Tokens
6,696
Response Time (avg)
6.84s
Total Cost
$0.280
Rank
#152
Total Output Tokens
487,519
Response Time (avg)
65.67s
Total Cost
$1.047
Recommended model GPT-6 Sol

It has the best score here (7.2), while costing about 3.7x less than Qwen3.5-122B-A10B (medium).

Detailed comparison

Metric GPT-6 Sol GPT-6 Sol none Release: 2026-09-23 Qwen3.5-122B-A10B Qwen3.5-122B-A10B medium Release: 2026-02-24
Score 7.2 7.1
Rank #148 #152
Reliability 10.0 10.0
Consistency 8.5 8.5
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 66.7% 71.2%
Flaky tests 4 4
Total Runs 66 66
Cost per result 2.151 8.446
Total Cost $0.280 $1.047
Input Price $2.000 / 1M $0.260 / 1M
Output Price $10.000 / 1M $2.080 / 1M
Total Input Tokens 106,332 124,780
Output Tokens 6,696 47,694
Reasoning Tokens 0 439,825
Response Time (avg) 6.84s 65.67s
Response Time (max) 23.89s 519.30s
Response Time (total) 150.43s 1444.68s
Parameters ~2T total (~150B active) 122B total (10B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#148 GPT-6 Sol

none
Cost
$0.030
Time
31.9s
Tokens
3,055 tok

#152 Qwen3.5-122B-A10B

medium
Cost
$0.019
Time
48.7s
Tokens
6,034 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-6 Sol 5.5 10.0 33.3% 0 12.23s 7,302 389 0
Qwen3.5-122B-A10B 6.0 7.2 55.6% 1 114.48s 7,630 8,057 82,578

Quick Compare

Switch Comparison Pair