Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 2.5 Flash vs MoonshotAI: Kimi K2.5

Last updated at: 2026-06-01

Metric Gemini 2.5 Flash Gemini 2.5 Flash none Release: 2025-06-17 Kimi K2.5 Kimi K2.5 medium Release: 2026-01-27
Score 6.4 6.7
Rank #95 #85
Reliability 10.0 10.0
Consistency 9.6 6.8
Tests Correct
Attempt pass rate 48.3% 66.7%
Flaky tests 1 8
Total Runs 60 60
Cost per result 0.159 3.486
Total Cost $0.015 $0.272
Input Price $0.300 / 1M $0.400 / 1M
Output Price $2.500 / 1M $1.900 / 1M
Output Tokens 1,764 48,374
Reasoning Tokens 0 128,473
Response Time (avg) 889ms 89.02s
Response Time (max) 4.39s 281.00s
Response Time (total) 17.79s 1157.32s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 582ms 102 0
Kimi K2.5 7.3 5.8 83.3% 2 51.38s 2,789 8,880
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 6.8 10.0 50.0% 0 810ms 477 0
Kimi K2.5 4.1 1.9 50.0% 2 215.89s 5,700 45,419
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 4.39s 366 0
Kimi K2.5 10.0 10.0 100.0% 0 71.37s 703 3,713
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 652ms 279 0
Kimi K2.5 10.0 10.0 100.0% 0 49.78s 563 7,940
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 5.9 7.2 55.6% 1 495ms 12 0
Kimi K2.5 3.5 4.4 33.3% 2 137.29s 20,753 30,564
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 5.0 10.0 0.0% 0 615ms 78 0
Kimi K2.5 6.5 3.4 66.7% 1 69.73s 3,815 4,262
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 590ms 72 0
Kimi K2.5 10.0 10.0 100.0% 0 92.47s 5,371 6,547
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 7.7 10.0 66.7% 0 604ms 132 0
Kimi K2.5 5.3 7.3 44.4% 1 43.23s 8,426 12,692
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 10.0 10.0 100.0% 0 1.91s 234 0
Kimi K2.5 10.0 10.0 100.0% 0 31.74s 242 812
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Gemini 2.5 Flash 3.0 10.0 0.0% 0 1.15s 12 0
Kimi K2.5 3.0 10.0 0.0% 0 83.95s 12 7,644

Quick Compare

Switch Comparison Pair