Navigate
AI BENCHY
Advertise here

Claude Sonnet 5 vs Gemma 4 31B (medium)

The average score is effectively tied at 6.1 vs 6.2. Gemma 4 31B (medium) has the lower benchmark cost at $0.101 vs $0.548. Claude Sonnet 5 is faster at 6.05s vs 78.59s, with pass rates of 40.9% vs 65.2%.

Last updated at: 2026-08-15

Rank
#166
Total Output Tokens
22,508
Response Time (avg)
6.05s
Total Cost
$0.548
Rank
#164
Total Output Tokens
267,421
Response Time (avg)
78.59s
Total Cost
$0.101
Recommended model Gemma 4 31B (medium)

It has the best score here (6.2), while costing about 5.4x less than Claude Sonnet 5.

Detailed comparison

Metric Claude Sonnet 5 Claude Sonnet 5 none Release: 2026-06-30 Gemma 4 31B Gemma 4 31B medium Release: 2026-04-02 Free Available
Score 6.1 6.2
Rank #166 #164
Reliability 10.0 10.0
Consistency 8.7 8.7
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 40.9% 65.2%
Flaky tests 4 3
Total Runs 66 66
Cost per result 7.817 1.152
Total Cost $0.548 $0.101
Input Price $2.000 / 1M $0.100 / 1M
Output Price $10.000 / 1M $0.340 / 1M
Total Input Tokens 161,032 94,969
Output Tokens 22,508 34,399
Reasoning Tokens 0 233,022
Response Time (avg) 6.05s 78.59s
Response Time (max) 33.39s 437.40s
Response Time (total) 132.99s 1571.84s
Parameters ~1T total (~100B active) 30.7B
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#166 Claude Sonnet 5

none
Cost
$0.061
Time
53.7s
Tokens
6,172 tok

#164 Gemma 4 31B

medium
Cost
$0.002
Time
45.7s
Tokens
2,696 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5 4.6 7.9 22.2% 1 3.67s 10,590 1,864 0
Gemma 4 31B 4.3 5.8 22.2% 1 219.76s 5,568 11,098 33,212

Quick Compare

Switch Comparison Pair