Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

DeepSeek: DeepSeek V4 Pro vs MoonshotAI: Kimi K2.5

Summary

DeepSeek V4 Pro vs Kimi K2.5 benchmark comparison: DeepSeek V4 Pro leads on average score with 6.3 vs 5.5. Kimi K2.5 has the lower benchmark cost at $0.028 vs $0.079. Kimi K2.5 is faster at 13.18s vs 65.21s, with pass rates of 52.4% vs 34.9%.

Recommended model: Kimi K2.5 - Its score stays close to the best score here (5.5 vs 6.3), while costing about 2.9x less than DeepSeek V4 Pro.

Last updated at: 2026-06-12

Metric DeepSeek V4 Pro DeepSeek V4 Pro high Release: 2026-04-24 Kimi K2.5 Kimi K2.5 none Release: 2026-01-27
Score 6.3 5.5
Rank #90 #121
Reliability 9.0 10.0
Consistency 7.6 8.9
Tests Correct
Attempt pass rate 52.4% 34.9%
Flaky tests 6 3
Total Runs 63 63
Cost per result 2.869 0.442
Total Cost $0.079 $0.028
Input Price $0.435 / 1M $0.400 / 1M
Output Price $0.870 / 1M $1.900 / 1M
Total Input Tokens 32,240 36,034
Output Tokens 12,250 6,657
Reasoning Tokens 72,257 0
Response Time (avg) 65.21s 13.18s
Response Time (max) 358.35s 42.13s
Response Time (total) 1304.19s 184.47s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#90 DeepSeek V4 Pro

high
Cost
$0.023
Time
257.6s
Tokens
14,870 tok

#121 MoonshotAI: Kimi K2.5

none
Cost
$0.015
Time
89.1s
Tokens
5,421 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 6.4 7.9 58.3% 1 16.53s 448 71 3,617
Kimi K2.5 3.6 8.4 8.3% 1 6.24s 652 373 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 3.3 6.4 11.1% 1 118.23s 1,966 111 20,940
Kimi K2.5 5.5 10.0 33.3% 0 24.56s 7,311 4,708 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 65.02s 14,016 465 5,914
Kimi K2.5 2.8 2.1 33.3% 1 19.16s 12,264 748 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 7.3 5.9 83.3% 1 23.62s 5,633 229 1,710
Kimi K2.5 7.3 5.8 83.3% 1 42.13s 7,180 187 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 2.9 7.2 11.1% 1 205.66s 430 10,529 28,089
Kimi K2.5 5.3 10.0 33.3% 0 4.38s 753 29 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 6.1 3.1 66.7% 1 25.09s 314 76 1,152
Kimi K2.5 10.0 10.0 100.0% 0 4.00s 483 76 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 41.16s 627 205 2,416
Kimi K2.5 6.5 10.0 50.0% 0 2.67s 677 60 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 5.9 7.2 55.6% 1 34.84s 544 139 4,019
Kimi K2.5 3.0 10.0 0.0% 0 4.04s 667 236 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 10.0 10.0 100.0% 0 21.33s 8,079 372 593
Kimi K2.5 10.0 10.0 100.0% 0 13.99s 5,835 220 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
DeepSeek V4 Pro 3.0 10.0 0.0% 0 39.14s 183 53 3,807
Kimi K2.5 3.0 10.0 0.0% 0 3.90s 212 20 0

Quick Compare

Switch Comparison Pair