Navigate
AI BENCHY
Advertise here

GPT-5.3 Chat vs Qwen3.8 2.4T A95B (high)

The average score is effectively tied at 7.5 vs 7.5. GPT-5.3 Chat has the lower benchmark cost at $0.571 vs $2.755. GPT-5.3 Chat is faster at 6.88s vs 134.41s, with pass rates of 68.2% vs 77.3%.

Last updated at: 2026-08-12

Rank
#81
Total Output Tokens
30,854
Response Time (avg)
6.88s
Total Cost
$0.571
Rank
#85
Total Output Tokens
512,351
Response Time (avg)
134.41s
Total Cost
$2.755
Recommended model GPT-5.3 Chat

It has the best score here (7.5), while costing about 4.8x less than Qwen3.8 2.4T A95B (high).

Detailed comparison

Metric GPT-5.3 Chat GPT-5.3 Chat none Release: 2026-03-03 Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B high Release: 2026-08-12
Score 7.5 7.5
Rank #81 #85
Reliability 10.0 8.8
Consistency 8.2 8.4
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 68.2% 77.3%
Flaky tests 5 4
Total Runs 66 66
Cost per result 4.387 18.365
Total Cost $0.571 $2.755
Input Price $1.750 / 1M $2.000 / 1M
Output Price $14.000 / 1M $6.000 / 1M
Total Input Tokens 78,990 109,036
Output Tokens 30,854 103,736
Reasoning Tokens 0 408,615
Response Time (avg) 6.88s 134.41s
Response Time (max) 18.33s 532.28s
Response Time (total) 151.31s 2956.97s
Parameters ~400B total (~17B active) 2.4T total (95B active)
Availability Closed Weights available

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#81 GPT-5.3 Chat

none
Cost
$0.008
Time
8.1s
Tokens
634 tok

#85 Qwen3.8 2.4T A95B

high
Invalid SVG
Cost
$0.000
Time
600.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.3 Chat 5.6 4.7 55.6% 2 10.52s 7,302 6,632 0
Qwen3.8 2.4T A95B 5.9 6.3 55.6% 1 160.52s 6,399 3,310 48,078

Quick Compare

Switch Comparison Pair