Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Ember-1 (high) vs GLM 5.1 (medium)

Ember-1 (high) leads on average score with 7.1 vs 7.0. GLM 5.1 (medium) has the lower benchmark cost at $1.053 vs $1.798. Ember-1 (high) is faster at 56.77s vs 58.80s, with pass rates of 71.2% vs 65.2%.

Last updated at: 2026-09-28

Compared models

Rank
#158
Total Output Tokens
110,125
Response Time (avg)
56.77s
Total Cost
$1.798
Rank
#164
Total Output Tokens
215,889
Response Time (avg)
58.80s
Total Cost
$1.053
Recommended model GLM 5.1 (medium)

Its score stays close to the best score here (7.0 vs 7.1), while costing about 1.7x less than Ember-1 (high).

Detailed comparison

Metric Ember-1 Ember-1 high Release: 2026-09-28 GLM 5.1 GLM 5.1 medium Release: 2026-04-07
Score 7.1 7.0
Rank #158 #164
Reliability 8.0 8.1
Consistency 7.3 8.4
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 71.2% 65.2%
Flaky tests 7 4
Total Runs 66 66
Cost per result 14.981 6.923
Total Cost $1.798 $1.053
Input Price $3.000 / 1M $1.400 / 1M
Output Price $15.000 / 1M $4.400 / 1M
Total Input Tokens 48,577 82,592
Output Tokens 7,345 16,089
Reasoning Tokens 102,780 199,800
Response Time (avg) 56.77s 58.80s
Response Time (max) 255.06s 308.75s
Response Time (total) 1248.88s 1234.72s
Parameters ~2.8T total (~104B active) 744B total (40B active)
Availability Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#158 Ember-1

high
Cost
$0.038
Time
55.1s
Tokens
2,671 tok

#164 GLM 5.1

medium
Reached the allocated time limit (300 seconds) without receiving showcase output.
Cost
$0.000
Time
300.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ember-1 8.2 7.2 88.9% 1 72.23s 7,144 1,281 23,093
GLM 5.1 4.6 3.7 44.5% 2 109.63s 5,702 4,871 37,826

Quick Compare

Switch Comparison Pair