Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

GPT-5.4 Mini (medium) vs Qwen3.8 2.4T A95B (high)

The average score is effectively tied at 7.5 vs 7.5. GPT-5.4 Mini (medium) has the lower benchmark cost at $0.756 vs $2.755. GPT-5.4 Mini (medium) is faster at 25.94s vs 134.41s, with pass rates of 71.2% vs 77.3%.

Last updated at: 2026-08-12

Rank
#83
Total Output Tokens
151,755
Response Time (avg)
25.94s
Total Cost
$0.756
Rank
#85
Total Output Tokens
512,351
Response Time (avg)
134.41s
Total Cost
$2.755
Recommended model GPT-5.4 Mini (medium)

It has the best score here (7.5), while costing about 3.6x less than Qwen3.8 2.4T A95B (high).

Detailed comparison

Metric GPT-5.4 Mini GPT-5.4 Mini medium Release: 2026-03-17 Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B high Release: 2026-08-12
Score 7.5 7.5
Rank #83 #85
Reliability 10.0 8.8
Consistency 7.7 8.4
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 71.2% 77.3%
Flaky tests 6 4
Total Runs 66 66
Cost per result 6.299 18.365
Total Cost $0.756 $2.755
Input Price $0.750 / 1M $2.000 / 1M
Output Price $4.500 / 1M $6.000 / 1M
Total Input Tokens 97,155 109,036
Output Tokens 6,211 103,736
Reasoning Tokens 145,544 408,615
Response Time (avg) 25.94s 134.41s
Response Time (max) 138.75s 532.28s
Response Time (total) 570.66s 2956.97s
Parameters ~400B total (~17B active) 2.4T total (95B active)
Availability Closed Weights available

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#83 GPT-5.4 Mini

medium
Cost
$0.056
Time
95.5s
Tokens
12,464 tok

#85 Qwen3.8 2.4T A95B

high
Invalid SVG
Cost
$0.000
Time
600.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 Mini 8.4 7.4 88.9% 1 57.87s 7,305 467 40,902
Qwen3.8 2.4T A95B 5.9 6.3 55.6% 1 160.52s 6,399 3,310 48,078

Quick Compare

Switch Comparison Pair