Navigate
AI BENCHY
Advertise here

Claude Opus 4.8 vs Gemma 4 26B A4B (medium)

Claude Opus 4.8 leads on average score with 7.3 vs 6.6. Gemma 4 26B A4B (medium) has the lower benchmark cost at $0.096 vs $1.166. Claude Opus 4.8 is faster at 4.91s vs 103.83s, with pass rates of 63.6% vs 66.7%.

Last updated at: 2026-07-25

Rank
#74
Total Output Tokens
16,797
Response Time (avg)
4.91s
Total Cost
$1.166
Rank
#104
Total Output Tokens
247,527
Response Time (avg)
103.83s
Total Cost
$0.096
Recommended model Claude Opus 4.8

It has the best score here (7.3), while responding about 21.1x faster than Gemma 4 26B A4B (medium).

Detailed comparison

Metric Claude Opus 4.8 Claude Opus 4.8 none Release: 2026-05-28 Gemma 4 26B A4B Gemma 4 26B A4B medium Release: 2026-04-03 Free Available
Score 7.3 6.6
Rank #74 #104
Reliability 10.0 9.4
Consistency 9.2 9.2
Tests Correct
Attempt pass rate 63.6% 66.7%
Flaky tests 2 2
Total Runs 66 66
Cost per result 8.969 0.643
Total Cost $1.166 $0.096
Input Price $5.000 / 1M $0.120 / 1M
Output Price $25.000 / 1M $0.350 / 1M
Total Input Tokens 149,206 77,550
Output Tokens 16,797 28,036
Reasoning Tokens 0 219,491
Response Time (avg) 4.91s 103.83s
Response Time (max) 35.03s 912.19s
Response Time (total) 108.03s 2180.47s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#74 Claude Opus 4.8

none
Cost
$0.053
Time
22.0s
Tokens
2,253 tok

#104 Gemma 4 26B A4B

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Opus 4.8 5.5 10.0 33.3% 0 3.29s 10,590 1,332 0
Gemma 4 26B A4B 2.9 10.0 0.0% 0 272.54s 5,062 14,838 44,567

Quick Compare

Switch Comparison Pair