Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Claude Sonnet 4.6 (medium) vs Claude Sonnet 5 (medium) vs Claude Opus 4.8 (medium) vs GLM 5.2 (medium) benchmark comparison: Claude Opus 4.8 (medium) leads on Score with 8.8. Claude Sonnet 4.6 (medium) leads on Reliability with 10.0. GLM 5.2 (medium) has the lowest Total Cost at $0.146. Claude Sonnet 5 (medium) is fastest at 12.54s.

Last updated at: 2026-08-14

Rank
#65
Total Output Tokens
120,427
Response Time (avg)
28.08s
Total Cost
$2.126
Rank
#46
Total Output Tokens
63,204
Response Time (avg)
12.54s
Total Cost
$0.922
Rank
#26
Total Output Tokens
51,927
Response Time (avg)
12.97s
Total Cost
$1.983
Rank
#63
Total Output Tokens
61,761
Response Time (avg)
23.28s
Total Cost
$0.146
Recommended model GLM 5.2 (medium)

It offers the best overall trade-off: a competitive score (7.8), lower cost than the other models in this comparison, and balanced response time.

Detailed comparison

Metric Claude Sonnet 4.6 Claude Sonnet 4.6 medium Release: 2026-02-17 Claude Sonnet 5 Claude Sonnet 5 medium Release: 2026-06-30 Claude Opus 4.8 Claude Opus 4.8 medium Release: 2026-05-28 GLM 5.2 GLM 5.2 medium Release: 2026-06-17
Score 7.8 8.2 8.8 7.8
Rank #65 #46 #26 #63
Reliability 10.0 10.0 10.0 9.5
Consistency 9.5 9.0 9.2 8.0
Attempts 66/66 66/66 66/66 63/66
Tests Correct
Attempt pass rate 65.2% 75.8% 83.3% 80.3%
Flaky tests 1 3 2 4
Total Runs 66 66 66 63
Cost per result 15.182 6.144 11.662 2.159
Total Cost $2.126 $0.922 $1.983 $0.146
Input Price $3.000 / 1M $2.000 / 1M $5.000 / 1M $0.630 / 1M
Output Price $15.000 / 1M $10.000 / 1M $25.000 / 1M $1.981 / 1M
Total Input Tokens 106,316 145,953 138,448 37,199
Output Tokens 79,338 52,330 40,765 12,261
Reasoning Tokens 41,089 10,874 11,162 49,500
Response Time (avg) 28.08s 12.54s 12.97s 23.28s
Response Time (max) 140.96s 66.71s 70.54s 101.36s
Response Time (total) 421.15s 275.90s 285.29s 488.94s
Parameters ~1T total (~100B active) ~1T total (~100B active) ~5T total (~500B active) 744B total (40B active)
Availability Closed Closed Closed Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#65 Claude Sonnet 4.6

medium
Invalid SVG
Cost
$0.000
Time
300.0s
Tokens
0 tok

#46 Claude Sonnet 5

medium
Cost
$0.007
Time
6.4s
Tokens
832 tok

#26 Claude Opus 4.8

medium
Cost
$0.057
Time
23.1s
Tokens
2,412 tok

#63 GLM 5.2

medium
Cost
$0.041
Time
195.8s
Tokens
9,287 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 4.6 5.7 6.6 44.4% 1 33.29s 6,995 16,089 3,686
Claude Sonnet 5 9.0 7.9 88.9% 1 17.28s 10,590 13,153 2,379
Claude Opus 4.8 10.0 10.0 100.0% 0 15.33s 10,590 9,945 1,381
GLM 5.2 8.2 7.2 88.9% 1 40.96s 7,317 1,475 17,123

Quick Compare

Switch Comparison Pair