Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Anthropic: Claude Opus 4.6 vs MoonshotAI: Kimi K2.6

Last updated at: 2026-05-19

Metric Claude Opus 4.6 Claude Opus 4.6 medium Release: 2026-02-05 Kimi K2.6 Kimi K2.6 medium Release: 2026-04-20
Score 7.4 7.6
Rank #57 #47
Reliability 10.0 10.0
Consistency 9.1 8.7
Tests Correct
Attempt pass rate 66.7% 71.9%
Flaky tests 2 3
Total Runs 57 57
Cost per result 14.243 6.476
Total Cost $1.710 $0.778
Input Price $5.000 / 1M $0.730 / 1M
Output Price $25.000 / 1M $3.490 / 1M
Output Tokens 37,874 96,469
Reasoning Tokens 21,390 195,991
Response Time (avg) 24.59s 49.92s
Response Time (max) 83.40s 215.85s
Response Time (total) 295.08s 898.64s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 6.4 5.8 66.7% 2 7.45s 986 1,071
Kimi K2.6 7.0 8.0 66.7% 1 11.59s 7,115 8,934
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 23.11s 3,486 1,504
Kimi K2.6 10.0 10.0 100.0% 0 106.96s 3,236 18,817
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 76.66s 8,178 5,194
Kimi K2.6 10.0 10.0 100.0% 0 40.96s 711 13,876
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 7.37s 691 757
Kimi K2.6 10.0 10.0 100.0% 0 20.38s 316 11,305
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 3.0 10.0 0.0% 0 83.40s 14,642 8,687
Kimi K2.6 5.3 7.2 44.4% 1 202.38s 47,035 98,262
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 5.04s 188 292
Kimi K2.6 10.0 10.0 100.0% 0 17.83s 3,981 4,472
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 2.43s 266 467
Kimi K2.6 10.0 10.0 100.0% 0 12.53s 3,977 5,269
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 7.7 10.0 66.7% 0 4.60s 531 637
Kimi K2.6 6.0 7.4 55.6% 1 25.59s 14,140 17,868
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 10.0 10.0 100.0% 0 9.73s 861 329
Kimi K2.6 10.0 10.0 100.0% 0 8.92s 248 1,011
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Claude Opus 4.6 3.0 10.0 0.0% 0 63.24s 8,045 2,452
Kimi K2.6 3.0 10.0 0.0% 0 130.27s 15,710 16,177

Quick Compare

Switch Comparison Pair