Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Qwen3.8 Max — low vs medium

medium leads on average score with 9.1 vs 8.7. low has the lower benchmark cost at $0.687 vs $0.728. low is faster at 21.19s vs 21.53s, with pass rates of 84.9% vs 84.9%.

Last updated at: 2026-08-04

Rank
#23
Total Output Tokens
81,411
Response Time (avg)
21.19s
Total Cost
$0.687
Rank
#14
Total Output Tokens
84,826
Response Time (avg)
21.53s
Total Cost
$0.728
Recommended model medium

It has the strongest score in this comparison (9.1) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric Qwen3.8 Max Qwen3.8 Max low Release: 2026-08-04 Qwen3.8 Max Qwen3.8 Max medium Release: 2026-08-04
Score 8.7 9.1
Rank #23 #14
Reliability 10.0 10.0
Consistency 9.0 9.7
Benchmark coverage 22/22 tests · 66/66 attempts 22/22 tests · 66/66 attempts
Tests Correct
Attempt pass rate 84.9% 84.9%
Flaky tests 3 1
Total Runs 66 66
Cost per result 4.041 4.040
Total Cost $0.687 $0.728
Input Price $2.000 / 1M $2.000 / 1M
Output Price $6.000 / 1M $6.000 / 1M
Total Input Tokens 99,172 109,046
Output Tokens 6,132 6,567
Reasoning Tokens 75,279 78,259
Response Time (avg) 21.19s 21.53s
Response Time (max) 155.75s 126.60s
Response Time (total) 466.14s 473.75s

Model generation showcase

Solar System Animation

Prompt: Create a single-file HTML/CSS animation of the Solar System against a dark background, using no JavaScript or external images. The output must feature a stationary Sun and all 8 primary planets (Mercury through Neptune) orbiting it continuously in the correct order using infinite CSS loops. Also include the Moon orbiting Earth and add labels to the planets. The design must be scalable and optimized for a 1:1 aspect ratio by default.

#23 Qwen3.8 Max

low
Cost
$0.042
Time
103.1s
Tokens
7,121 tok

#14 Qwen3.8 Max

medium
Cost
$0.104
Time
286.9s
Tokens
17,399 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Qwen3.8 Max 8.4 7.4 88.9% 1 28.69s 8,127 512 15,952
Qwen3.8 Max 10.0 10.0 100.0% 0 41.35s 7,893 430 23,804

Quick Compare

Switch Comparison Pair