Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Google: Gemini 3.5 Flash-Lite vs Kwaipilot: KAT-Coder-Pro V2.5

The average score is effectively tied at 6.7 vs 6.7. Gemini 3.5 Flash-Lite (low) has the lower benchmark cost at $0.145 vs $0.476. Gemini 3.5 Flash-Lite (low) is faster at 2.25s vs 25.56s, with pass rates of 66.7% vs 68.2%.

Recommended modelGemini 3.5 Flash-Lite (low)It has the best score here (6.7), while costing about 3.3x less than KAT-Coder-Pro V2.5.

Last updated at: 2026-07-21

Metric Gemini 3.5 Flash-Lite Gemini 3.5 Flash-Lite low Release: 2026-07-21 KAT-Coder-Pro V2.5 KAT-Coder-Pro V2.5 none Release: 2026-07-14
Score 6.7 6.7
Rank #95 #97
Reliability 10.0 10.0
Consistency 8.2 7.4
Tests Correct
Attempt pass rate 66.7% 68.2%
Flaky tests 5 7
Total Runs 66 66
Cost per result 1.201 4.319
Total Cost $0.145 $0.476
Input Price $0.300 / 1M $0.740 / 1M
Output Price $2.500 / 1M $2.960 / 1M
Total Input Tokens 144,622 98,499
Output Tokens 15,302 135,861
Reasoning Tokens 24,971 0
Response Time (avg) 2.25s 25.56s
Response Time (max) 13.50s 335.41s
Response Time (total) 49.58s 562.43s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#95 Gemini 3.5 Flash-Lite

low
Cost
$0.008
Time
7.7s
Tokens
3,056 tok

#97 KAT-Coder-Pro V2.5

none
Cost
$0.010
Time
29.0s
Tokens
3,439 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Gemini 3.5 Flash-Lite 4.1 7.9 11.1% 1 672ms 8,136 463 0
KAT-Coder-Pro V2.5 6.1 4.7 66.7% 2 22.52s 7,893 22,440 0

Quick Compare

Switch Comparison Pair