Navigate
AI BENCHY
Advertise here

GPT-4o-mini vs Hy4 preview

GPT-4o-mini leads on average score with 5.0 vs 4.8. GPT-4o-mini has the lower benchmark cost at $0.010 vs $0.041. GPT-4o-mini is faster at 1.92s vs 39.24s, with pass rates of 22.7% vs 16.7%.

Last updated at: 2026-08-29

Rank
#249
Total Output Tokens
2,913
Response Time (avg)
1.92s
Total Cost
$0.010
Rank
#253
Total Output Tokens
7,306
Response Time (avg)
39.24s
Total Cost
$0.041
Recommended model GPT-4o-mini

It has the best score here (5.0), while costing about 4.1x less than Hy4 preview.

Detailed comparison

Metric GPT-4o-mini GPT-4o-mini none Release: 2024-07-18 Hy4 preview Hy4 preview none Release: 2026-08-29
Score 5.0 4.8
Rank #249 #253
Reliability 10.0 4.3
Consistency 9.9 9.3
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 22.7% 16.7%
Flaky tests 0 1
Total Runs 66 66
Cost per result 0.195 1.343
Total Cost $0.010 $0.041
Input Price $0.150 / 1M $0.834 / 1M
Output Price $0.600 / 1M $2.501 / 1M
Total Input Tokens 53,145 26,376
Output Tokens 2,913 7,306
Reasoning Tokens 0 0
Response Time (avg) 1.92s 39.24s
Response Time (max) 7.58s 95.32s
Response Time (total) 30.71s 863.30s
Parameters ~8B 770B total (49B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#249 GPT-4o-mini

none
Cost
$0.001
Time
6.6s
Tokens
742 tok

#253 Hy4 preview

none
This operation was aborted
Cost
$0.000
Time
0.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-4o-mini 3.2 9.6 0.0% 0 1.63s 7,314 367 0
Hy4 preview 3.8 9.4 0.0% 0 45.53s 6,646 4,103 0

Quick Compare

Switch Comparison Pair