Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Gemini 3.1 Pro Preview (medium) vs Qwen3.8 Max (0902) (low)

Gemini 3.1 Pro Preview (medium) leads on average score with 9.2 vs 9.1. Qwen3.8 Max (0902) (low) has the lower benchmark cost at $0.664 vs $1.352. Gemini 3.1 Pro Preview (medium) is faster at 20.62s vs 24.21s, with pass rates of 90.9% vs 86.4%.

Last updated at: 2026-09-07

Rank
#18
Total Output Tokens
97,238
Response Time (avg)
20.62s
Total Cost
$1.352
Rank
#23
Total Output Tokens
75,865
Response Time (avg)
24.21s
Total Cost
$0.664
Recommended model Qwen3.8 Max (0902) (low)

Its score stays close to the best score here (9.1 vs 9.2), while costing about 2.0x less than Gemini 3.1 Pro Preview (medium).

Detailed comparison

Metric Gemini 3.1 Pro Preview Gemini 3.1 Pro Preview medium Release: 2026-02-19 Qwen3.8 Max (0902) Qwen3.8 Max (0902) low Release: 2026-09-07
Score 9.2 9.1
Rank #18 #23
Reliability 10.0 10.0
Consistency 10.0 9.3
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 90.9% 86.4%
Flaky tests 0 2
Total Runs 66 66
Cost per result 6.758 3.684
Total Cost $1.352 $0.664
Input Price $2.000 / 1M $2.000 / 1M
Output Price $12.000 / 1M $6.000 / 1M
Total Input Tokens 92,296 103,954
Output Tokens 5,232 6,322
Reasoning Tokens 92,006 69,543
Response Time (avg) 20.62s 24.21s
Response Time (max) 88.68s 207.79s
Response Time (total) 329.94s 532.62s
Parameters ~1.2T total (~20B active) 2.4T total (~100B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#18 Gemini 3.1 Pro Preview

medium
Cost
$0.115
Time
87.2s
Tokens
9,629 tok

#23 Qwen3.8 Max (0902)

low
Cost
$0.014
Time
41.2s
Tokens
2,421 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.1 Pro Preview 7.9 9.9 66.7% 0 40.17s 8,124 435 41,247
Qwen3.8 Max (0902) 10.0 10.0 100.0% 0 37.42s 8,127 485 17,404

Quick Compare

Switch Comparison Pair