Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

GPT-6 Astra (high) vs Gemini 3.8 Flash (high) vs Claude Fable 5.1 (high) benchmark comparison: Gemini 3.8 Flash (high) leads on Score with 9.9. Gemini 3.8 Flash (high) leads on Reliability with 10.0. Gemini 3.8 Flash (high) has the lowest Total Cost at $1.595. Claude Fable 5.1 (high) is fastest at 13.59s.

Last updated at: 2026-09-04

Rank
#4
Total Output Tokens
53,622
Response Time (avg)
18.65s
Total Cost
$3.474
Rank
#2
Total Output Tokens
408,188
Response Time (avg)
37.19s
Total Cost
$1.595
Rank
#25
Total Output Tokens
40,071
Response Time (avg)
13.59s
Total Cost
$3.364
Recommended model Gemini 3.8 Flash (high)

It has the best score here (9.9), while costing about 2.1x less than the other models in this comparison.

Detailed comparison

Metric GPT-6 Astra GPT-6 Astra high Release: 2026-09-04 Gemini 3.8 Flash Gemini 3.8 Flash high Release: 2026-09-02 Claude Fable 5.1 Claude Fable 5.1 high Release: 2026-09-02
Score 9.9 9.9 9.0
Rank #4 #2 #25
Reliability 9.9 10.0 10.0
Consistency 9.6 9.6 9.6
Attempts 66/66 66/66 66/66
Tests Correct
Attempt pass rate 97.0% 98.5% 89.4%
Flaky tests 1 1 1
Total Runs 66 66 66
Cost per result 16.543 7.592 17.701
Total Cost $3.474 $1.595 $3.364
Input Price $10.000 / 1M $0.750 / 1M $10.000 / 1M
Output Price $50.000 / 1M $3.750 / 1M $50.000 / 1M
Total Input Tokens 79,290 84,593 135,958
Output Tokens 4,847 10,517 7,930
Reasoning Tokens 48,775 397,671 32,141
Response Time (avg) 18.65s 37.19s 13.59s
Response Time (max) 254.93s 177.08s 40.17s
Response Time (total) 410.28s 818.19s 298.90s
Parameters ~2T total (~150B active) ~500B total (~20B active) ~9.5T total (~878B active)
Availability Closed Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#4 GPT-6 Astra

high
Cost
$0.457
Time
176.7s
Tokens
9,218 tok

#2 Gemini 3.8 Flash

high
Cost
$0.077
Time
129.1s
Tokens
20,373 tok

#25 Claude Fable 5.1

high
Cost
$0.384
Time
102.5s
Tokens
7,824 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-6 Astra 10.0 10.0 100.0% 0 5.86s 7,302 392 2,387
Gemini 3.8 Flash 10.0 10.0 100.0% 0 29.30s 8,118 462 50,196
Claude Fable 5.1 8.4 7.4 88.9% 1 13.91s 10,608 520 6,503

Quick Compare

Switch Comparison Pair