Navigate
Advertise here

Granite 4.2 8B vs Mercury 2.5

Granite 4.2 8B leads on average score with 4.3 vs 3.9. Mercury 2.5 has the lower benchmark cost at $0.016 vs $0.026. Mercury 2.5 is faster at 2.45s vs 86.27s, with pass rates of 16.7% vs 25.8%.

Last updated at: 2026-09-08

Compared models

Rank
#302
Total Output Tokens
130,805
Response Time (avg)
86.27s
Total Cost
$0.026
Rank
#312
Total Output Tokens
71,143
Response Time (avg)
2.45s
Total Cost
$0.016
Recommended model Mercury 2.5

Its score stays close to the best score here (3.9 vs 4.3), while costing about 1.6x less than Granite 4.2 8B.

Detailed comparison

Metric Granite 4.2 8B Granite 4.2 8B none Release: 2026-09-02 Mercury 2.5 Mercury 2.5 none Release: 2026-09-08
Score 4.3 3.9
Rank #302 #312
Reliability 8.6 10.0
Consistency 9.3 7.5
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 16.7% 25.8%
Flaky tests 2 7
Total Runs 66 66
Cost per result 0.847 0.782
Total Cost $0.026 $0.016
Input Price $0.100 / 1M $0.040 / 1M
Output Price $0.150 / 1M $0.150 / 1M
Total Input Tokens 57,786 123,740
Output Tokens 130,805 71,143
Reasoning Tokens 0 0
Response Time (avg) 86.27s 2.45s
Response Time (max) 747.17s 36.48s
Response Time (total) 1897.96s 53.83s
Parameters 8B ~100B
Availability Open source Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#302 IBM: Granite 4.2 8B

none
Provider returned error
Cost
$0.000
Time
0.2s
Tokens
0 tok

#312 Mercury 2.5

none
Cost
$0.001
Time
2.7s
Tokens
2,749 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Granite 4.2 8B 5.5 10.0 33.3% 0 105.60s 8,358 47,944 0
Mercury 2.5 3.7 7.0 22.2% 1 12.50s 7,888 61,297 0

Quick Compare

Switch Comparison Pair