Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

MoonshotAI: Kimi K2.5 vs Tencent: Hy3 preview

Last updated at: 2026-05-22

Metric Kimi K2.5 Kimi K2.5 medium Release: 2026-01-27 Hy3 preview Hy3 preview high Release: 2026-04-22
Score 6.7 8.0
Rank #79 #22
Reliability 10.0 10.0
Consistency 6.8 9.5
Tests Correct
Attempt pass rate 66.7% 77.1%
Flaky tests 8 1
Total Runs 60 60
Cost per result 3.479 0.000
Total Cost $0.314 $0.000
Input Price $0.400 / 1M $0.066 / 1M
Output Price $1.900 / 1M $0.260 / 1M
Output Tokens 46,619 216,503
Reasoning Tokens 128,184 0
Response Time (avg) 89.36s 56.77s
Response Time (max) 281.00s 149.94s
Response Time (total) 1161.65s 851.49s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 7.3 5.8 83.3% 2 51.38s 2,789 8,880
Hy3 preview 8.9 10.0 100.0% 0 15.12s 6,839 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 4.1 1.9 50.0% 2 215.89s 5,700 45,419
Hy3 preview 10.0 10.0 100.0% 0 99.76s 38,167 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 10.0 10.0 100.0% 0 71.37s 703 3,713
Hy3 preview 10.0 10.0 100.0% 0 113.09s 31,319 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 10.0 10.0 100.0% 0 49.78s 563 7,940
Hy3 preview 6.5 10.0 50.0% 0 12.11s 4,323 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 3.5 4.4 33.3% 2 137.29s 20,753 30,564
Hy3 preview 5.3 7.2 44.4% 1 109.04s 87,559 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 6.5 3.4 66.7% 1 69.73s 3,815 4,262
Hy3 preview 0.0 0.0 0.0% 0 0ms 0 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 10.0 10.0 100.0% 0 92.47s 5,371 6,547
Hy3 preview 9.9 10.0 100.0% 0 34.02s 13,331 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 5.3 7.3 44.4% 1 45.40s 6,671 12,403
Hy3 preview 10.0 10.0 100.0% 0 29.74s 15,503 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 10.0 10.0 100.0% 0 31.74s 242 812
Hy3 preview 10.0 10.0 100.0% 0 78.83s 10,370 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Kimi K2.5 3.0 10.0 0.0% 0 83.95s 12 7,644
Hy3 preview 3.0 10.0 0.0% 0 47.71s 9,092 0

Quick Compare

Switch Comparison Pair