Navigate
Advertise here

GPT-5.4 (medium) vs Qwen3.8 27B (high)

GPT-5.4 (medium) leads on average score with 8.8 vs 8.7. Qwen3.8 27B (high) has the lower benchmark cost at ~$0.298 vs $2.006. GPT-5.4 (medium) is faster at 24.81s vs 166.34s, with pass rates of 78.3% vs 78.3%.

Last updated at: 2026-10-01

Compared models

Rank
#50
Total Output Tokens
96,502
Response Time (avg)
24.81s
Total Cost
$2.006
Rank
#55
Total Output Tokens
677,080
Response Time (avg)
166.34s
Total Cost
~$0.298
Recommended model Qwen3.8 27B (high)

Its score stays close to the best score here (8.7 vs 8.8), while costing about 6.7x less than GPT-5.4 (medium).

Detailed comparison

Metric GPT-5.4 GPT-5.4 medium Release: 2026-03-05 Qwen3.8 27B Qwen3.8 27B high Release: 2026-08-14 Free Available
Score 8.8 8.7
Rank #50 #55
Reliability 10.0 9.7
Consistency 8.7 9.2
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 78.3% 78.3%
Flaky tests 4 2
Total Runs 69 69
Cost per result 12.535 ~1.751
Total Cost $2.006 ~$0.298
Input Price $2.500 / 1M N/A
Output Price $15.000 / 1M N/A
Total Input Tokens 223,210 348,967
Output Tokens 7,312 1,005
Reasoning Tokens 89,190 676,075
Response Time (avg) 24.81s 166.34s
Response Time (max) 100.41s 640.85s
Response Time (total) 570.60s 3825.81s
Parameters ~1.5T total (~100B active) 27.3B
Availability Closed Weights available

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#50 GPT-5.4

medium
Cost
$0.214
Time
199.6s
Tokens
14,349 tok

#55 Qwen3.8 27B

high
Cost
~$0.008
Time
270.1s
Tokens
18,963 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 8.8 7.8 88.9% 1 44.36s 7,305 433 24,216
Qwen3.8 27B 8.1 7.0 88.9% 1 248.62s 8,235 346 143,335

Quick Compare

Switch Comparison Pair