Navigate
AI BENCHY
Compare Charts Methodology
❤️ Made by XCS
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Google: Gemini 3.1 Flash Lite Preview vs MoonshotAI: Kimi K2.5

Compare:

Last updated at: 2026-03-05

Metric Google: Gemini 3.1 Flash Lite Preview low Release: 2026-03-03 MoonshotAI: Kimi K2.5 medium Release: 2026-01-27
Avg Score 7.6 6.4
Rank #12 #29
Tests Correct
Consistency 10.0 7.8
Cost per result 0.170 2.082
Total Cost $0.019 $0.188
Attempt pass rate 73.3% 73.3%
Flaky tests 0 4
common.totalRuns 45 (15 x 3) 45 (15 x 3)
Output Tokens 1,542 34,638
Reasoning Tokens 6,888 68,234
Response Time (avg) 3.49s 69.84s
Response Time (max) 11.91s 137.29s
Response Time (total) 52.29s 558.72s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Avg Score vs Response Time (avg)

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 7.0 10.0 66.7% 0 2.18s 456 1,224
MoonshotAI: Kimi K2.5 7.0 7.2 88.9% 1 85.28s 335 6,255
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 10.0 10.0 0.0% 0 11.91s 225 762
MoonshotAI: Kimi K2.5 10.0 10.0 100.0% 0 71.37s 703 3,713
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 9.9 10.0 100.0% 0 3.00s 291 696
MoonshotAI: Kimi K2.5 9.9 10.0 100.0% 0 49.78s 563 7,940
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 4.0 10.0 33.3% 0 2.36s 18 1,212
MoonshotAI: Kimi K2.5 10.0 4.4 33.3% 2 137.29s 20,753 30,564
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 1.49s 72 753
MoonshotAI: Kimi K2.5 10.0 10.0 100.0% 0 92.47s 5,371 6,547
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 2.76s 243 1,248
MoonshotAI: Kimi K2.5 4.0 7.3 44.4% 1 45.40s 6,671 12,403
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Google: Gemini 3.1 Flash Lite Preview 10.0 10.0 100.0% 0 9.54s 237 993
MoonshotAI: Kimi K2.5 10.0 10.0 100.0% 0 31.74s 242 812

Quick Compare

Switch Comparison Pair