Navigate
Advertise here

Granite 4.2 8B (low) vs Mercury 2.5 (low)

The average score is effectively tied at 5.1 vs 5.1. Mercury 2.5 (low) has the lower benchmark cost at $0.011 vs $0.012. Mercury 2.5 (low) is faster at 1.35s vs 54.36s, with pass rates of 31.8% vs 39.4%.

Last updated at: 2026-09-08

Compared models

Rank
#263
Total Output Tokens
27,928
Response Time (avg)
54.36s
Total Cost
$0.012
Rank
#265
Total Output Tokens
33,719
Response Time (avg)
1.35s
Total Cost
$0.011
Recommended model Mercury 2.5 (low)

It has the best score here (5.1), while responding about 40.2x faster than Granite 4.2 8B (low).

Detailed comparison

Metric Granite 4.2 8B Granite 4.2 8B low Release: 2026-09-02 Mercury 2.5 Mercury 2.5 low Release: 2026-09-08
Score 5.1 5.1
Rank #263 #265
Reliability 9.0 9.8
Consistency 10.0 8.1
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 31.8% 39.4%
Flaky tests 0 5
Total Runs 66 66
Cost per result 0.167 0.177
Total Cost $0.012 $0.011
Input Price $0.100 / 1M $0.040 / 1M
Output Price $0.150 / 1M $0.150 / 1M
Total Input Tokens 74,721 138,020
Output Tokens 9,205 6,083
Reasoning Tokens 18,723 27,636
Response Time (avg) 54.36s 1.35s
Response Time (max) 566.80s 7.48s
Response Time (total) 1195.97s 29.74s
Parameters 8B ~100B
Availability Open source Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#263 IBM: Granite 4.2 8B

low
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok

#265 Mercury 2.5

low
Cost
$0.001
Time
2.4s
Tokens
1,236 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.2 8B 5.5 10.0 33.3% 0 13.01s 8,439 2,469 2,984
Mercury 2.5 5.5 10.0 33.3% 0 972ms 7,909 521 2,269

Quick Compare

Switch Comparison Pair