Navigate
Advertise here

Ember-1 (high) vs Mercury 2 (medium)

Ember-1 (high) leads on average score with 7.1 vs 7.0. Mercury 2 (medium) has the lower benchmark cost at $0.094 vs $1.798. Mercury 2 (medium) is faster at 2.95s vs 56.77s, with pass rates of 71.2% vs 53.0%.

Last updated at: 2026-09-28

Compared models

Rank
#158
Total Output Tokens
110,125
Response Time (avg)
56.77s
Total Cost
$1.798
Rank
#166
Total Output Tokens
87,554
Response Time (avg)
2.95s
Total Cost
$0.094
Recommended model Mercury 2 (medium)

Its score stays close to the best score here (7.0 vs 7.1), while costing about 19.3x less than Ember-1 (high).

Detailed comparison

Metric Ember-1 Ember-1 high Release: 2026-09-28 Mercury 2 Mercury 2 medium Release: 2026-02-24
Score 7.1 7.0
Rank #158 #166
Reliability 8.0 10.0
Consistency 7.3 8.4
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 71.2% 53.0%
Flaky tests 7 4
Total Runs 66 66
Cost per result 14.981 0.933
Total Cost $1.798 $0.094
Input Price $3.000 / 1M $0.250 / 1M
Output Price $15.000 / 1M $0.750 / 1M
Total Input Tokens 48,577 110,515
Output Tokens 7,345 10,325
Reasoning Tokens 102,780 77,229
Response Time (avg) 56.77s 2.95s
Response Time (max) 255.06s 14.63s
Response Time (total) 1248.88s 61.89s
Parameters ~2.8T total (~104B active) ~100B
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#158 Ember-1

high
Cost
$0.038
Time
55.1s
Tokens
2,671 tok

#166 Mercury 2

medium
Cost
$0.002
Time
2.1s
Tokens
1,702 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ember-1 8.2 7.2 88.9% 1 72.23s 7,144 1,281 23,093
Mercury 2 8.2 7.7 77.8% 1 2.04s 7,065 296 11,328

Quick Compare

Switch Comparison Pair