Navigate
Advertise here

Ember-1 (low) vs GPT-5.2 Chat

The average score is effectively tied at 7.0 vs 7.0. GPT-5.2 Chat has the lower benchmark cost at $0.594 vs $1.488. GPT-5.2 Chat is faster at 7.43s vs 34.90s, with pass rates of 65.2% vs 66.7%.

Last updated at: 2026-10-01

Compared models

Rank
#157
Total Output Tokens
56,766
Response Time (avg)
34.90s
Total Cost
$1.488
Rank
#153
Total Output Tokens
29,742
Response Time (avg)
7.43s
Total Cost
$0.594
Recommended model GPT-5.2 Chat

It has the best score here (7.0), while costing about 2.5x less than Ember-1 (low).

Detailed comparison

Metric Ember-1 Ember-1 low Release: 2026-09-28 GPT-5.2 Chat GPT-5.2 Chat none Release: 2025-12-11
Score 7.0 7.0
Rank #157 #153
Reliability 8.8 10.0
Consistency 8.2 8.7
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 65.2% 66.7%
Flaky tests 5 4
Total Runs 69 69
Cost per result 11.441 4.564
Total Cost $1.488 $0.594
Input Price $3.000 / 1M $1.750 / 1M
Output Price $15.000 / 1M $14.000 / 1M
Total Input Tokens 211,908 101,104
Output Tokens 14,158 29,742
Reasoning Tokens 42,608 0
Response Time (avg) 34.90s 7.43s
Response Time (max) 121.29s 38.52s
Response Time (total) 802.79s 163.35s
Parameters ~2.8T total (~104B active) ~400B total (~17B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#157 Ember-1

low
Cost
$0.016
Time
24.2s
Tokens
1,165 tok

#153 GPT-5.2 Chat

none
Cost
$0.010
Time
15.3s
Tokens
797 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Ember-1 10.0 10.0 100.0% 0 37.91s 8,097 2,518 13,142
GPT-5.2 Chat 8.8 7.8 88.9% 1 9.82s 7,305 6,731 0

Quick Compare

Switch Comparison Pair