Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

MoonshotAI: Kimi K2.6 vs OpenAI: GPT-5.3 Chat

Summary

Kimi K2.6 vs GPT-5.3 Chat benchmark comparison: The average score is effectively tied at 7.2 vs 7.2. GPT-5.3 Chat has the lower benchmark cost at $0.433 vs $0.891. GPT-5.3 Chat is faster at 6.34s vs 71.67s, with pass rates of 65.1% vs 66.7%.

Recommended model: GPT-5.3 Chat - It has the best score here (7.2), while costing about 2.1x less than Kimi K2.6.

Last updated at: 2026-06-04

Metric Kimi K2.6 Kimi K2.6 medium Release: 2026-04-20 Free Available GPT-5.3 Chat GPT-5.3 Chat none Release: 2026-03-03
Score 7.2 7.2
Rank #60 #63
Reliability 10.0 10.0
Consistency 8.6 8.1
Tests Correct
Attempt pass rate 65.1% 66.7%
Flaky tests 3 5
Total Runs 63 63
Cost per result 8.358 3.605
Total Cost $0.891 $0.433
Input Price $0.684 / 1M $1.750 / 1M
Output Price $3.420 / 1M $14.000 / 1M
Total Input Tokens 29,450 34,209
Output Tokens 102,923 26,617
Reasoning Tokens 254,094 0
Response Time (avg) 71.67s 6.34s
Response Time (max) 406.78s 18.33s
Response Time (total) 1433.36s 133.13s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#60 MoonshotAI: Kimi K2.6

medium
Cost
$0.013
Time
103.4s
Tokens
3,620 tok

#63 GPT-5.3 Chat

none
Cost
$0.008
Time
8.1s
Tokens
634 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 7.0 8.0 66.7% 1 11.59s 618 7,115 8,934
GPT-5.3 Chat 6.7 8.1 58.3% 1 3.86s 606 3,167 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 5.7 8.6 33.3% 0 214.42s 2,925 9,970 77,189
GPT-5.3 Chat 5.6 4.7 55.6% 2 10.52s 7,302 6,632 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 10.0 10.0 100.0% 0 40.96s 11,271 711 13,876
GPT-5.3 Chat 10.0 10.0 100.0% 0 11.96s 11,019 2,614 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 10.0 10.0 100.0% 0 20.38s 7,014 316 11,305
GPT-5.3 Chat 10.0 10.0 100.0% 0 2.21s 7,140 942 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 5.3 7.2 44.4% 1 202.38s 326 47,035 98,262
GPT-5.3 Chat 3.5 4.4 33.3% 2 13.01s 723 8,264 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 10.0 10.0 100.0% 0 17.83s 477 3,981 4,472
GPT-5.3 Chat 4.6 10.0 0.0% 0 1.99s 477 319 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 10.0 10.0 100.0% 0 12.53s 669 3,977 5,269
GPT-5.3 Chat 9.8 10.0 100.0% 0 3.51s 660 1,491 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 6.0 7.4 55.6% 1 25.06s 651 13,860 17,599
GPT-5.3 Chat 10.0 10.0 100.0% 0 2.99s 642 1,758 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 10.0 10.0 100.0% 0 8.92s 5,286 248 1,011
GPT-5.3 Chat 10.0 10.0 100.0% 0 8.36s 5,445 861 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.6 3.0 10.0 0.0% 0 130.27s 213 15,710 16,177
GPT-5.3 Chat 3.0 10.0 0.0% 0 4.38s 195 569 0

Quick Compare

Switch Comparison Pair