Navigate
Advertise here

Claude Sonnet 4.6 (medium) vs Qwen3.7 Flash (high)

Claude Sonnet 4.6 (medium) leads on average score with 7.8 vs 7.8. Qwen3.7 Flash (high) has the lower benchmark cost at $0.052 vs $2.126. Claude Sonnet 4.6 (medium) is faster at 28.08s vs 41.37s, with pass rates of 65.2% vs 71.2%.

Last updated at: 2026-09-10

Compared models

Rank
#87
Total Output Tokens
120,427
Response Time (avg)
28.08s
Total Cost
$2.126
Rank
#91
Total Output Tokens
370,250
Response Time (avg)
41.37s
Total Cost
$0.052
Recommended model Qwen3.7 Flash (high)

Its score stays close to the best score here (7.8 vs 7.8), while costing about 41.2x less than Claude Sonnet 4.6 (medium).

Detailed comparison

Metric Claude Sonnet 4.6 Claude Sonnet 4.6 medium Release: 2026-02-17 Qwen3.7 Flash Qwen3.7 Flash high Release: 2026-07-28
Score 7.8 7.8
Rank #87 #91
Reliability 10.0 10.0
Consistency 9.5 7.8
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 65.2% 71.2%
Flaky tests 1 6
Total Runs 66 66
Cost per result 15.182 0.398
Total Cost $2.126 $0.052
Input Price $3.000 / 1M $0.030 / 1M
Output Price $15.000 / 1M $0.130 / 1M
Total Input Tokens 106,316 116,331
Output Tokens 79,338 12,330
Reasoning Tokens 41,089 357,920
Response Time (avg) 28.08s 41.37s
Response Time (max) 140.96s 431.47s
Response Time (total) 421.15s 910.11s
Parameters ~1T total (~100B active) ~35B total (~3B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#87 Claude Sonnet 4.6

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

#91 Qwen3.7 Flash

high
Reached the allocated time limit (600 seconds) without receiving showcase output.
Cost
$0.000
Time
600.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 4.6 5.7 6.6 44.4% 1 33.29s 6,995 16,089 3,686
Qwen3.7 Flash 8.2 9.7 66.7% 0 58.84s 7,893 503 62,350

Quick Compare

Switch Comparison Pair