Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Kimi K3 (max) vs Kimi K2.6 (medium) vs GPT-5.6 Sol (medium) vs GPT-5.6 Terra (high) benchmark comparison: GPT-5.6 Sol (medium) leads on Score with 9.4. Kimi K2.6 (medium) leads on Reliability with 10.0. GPT-5.6 Sol (medium) has the lowest Total Cost at $0.508. GPT-5.6 Terra (high) is fastest at 11.22s.

Last updated at: 2026-08-29

Rank
#59
Total Output Tokens
224,128
Response Time (avg)
142.84s
Total Cost
$3.467
Rank
#107
Total Output Tokens
347,519
Response Time (avg)
115.89s
Total Cost
$1.247
Rank
#8
Total Output Tokens
34,977
Response Time (avg)
12.63s
Total Cost
$0.508
Rank
#58
Total Output Tokens
56,088
Response Time (avg)
11.22s
Total Cost
$0.836
Recommended model GPT-5.6 Sol (medium)

It has the best score here (9.4), while costing about 3.6x less than the other models in this comparison.

Detailed comparison

Metric Kimi K3 Kimi K3 max Release: 2026-07-16 Kimi K2.6 Kimi K2.6 medium Release: 2026-04-20 GPT-5.6 Sol GPT-5.6 Sol medium Release: 2026-07-09 GPT-5.6 Terra GPT-5.6 Terra high Release: 2026-07-09
Score 7.9 7.2 9.4 8.0
Rank #59 #107 #8 #58
Reliability 9.0 10.0 10.0 10.0
Consistency 9.2 8.3 8.9 9.3
Attempts 66/66 66/66 66/66 66/66
Tests Correct
Attempt pass rate 77.3% 63.6% 90.9% 68.2%
Flaky tests 2 4 3 2
Total Runs 66 66 66 66
Cost per result 21.668 10.029 8.025 7.293
Total Cost $3.467 $1.247 $0.508 $0.836
Input Price $3.000 / 1M $0.950 / 1M $2.000 / 1M $2.000 / 1M
Output Price $15.000 / 1M $4.000 / 1M $10.000 / 1M $12.000 / 1M
Total Input Tokens 34,960 68,913 79,006 81,056
Output Tokens 2,887 64,654 4,696 5,055
Reasoning Tokens 221,241 282,865 30,281 51,033
Response Time (avg) 142.84s 115.89s 12.63s 11.22s
Response Time (max) 766.58s 876.20s 79.40s 91.49s
Response Time (total) 2714.02s 2433.66s 277.88s 246.77s
Parameters 2.8T total (104B active) 1T total (32B active) ~2T total (~150B active) ~1T total (~60B active)
Availability Weights available Weights available Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#59 MoonshotAI: Kimi K3

max
Cost
$0.281
Time
506.9s
Tokens
18,863 tok

#107 MoonshotAI: Kimi K2.6

medium
Cost
$0.013
Time
103.4s
Tokens
3,620 tok

#8 GPT-5.6 Sol

medium
Cost
$0.124
Time
57.2s
Tokens
4,199 tok

#58 GPT-5.6 Terra

high
Cost
$0.055
Time
30.2s
Tokens
3,748 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K3 10.0 10.0 100.0% 0 325.79s 8,061 460 84,012
Kimi K2.6 5.7 8.6 33.3% 0 214.42s 2,925 9,970 77,189
GPT-5.6 Sol 10.0 10.0 100.0% 0 9.40s 7,302 394 4,706
GPT-5.6 Terra 7.6 7.2 77.8% 1 9.14s 7,302 370 6,087

Quick Compare

Switch Comparison Pair