Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Kimi K2.5 vs Qwen3.6 27B

The average score is effectively tied at 5.5 vs 5.5. Qwen3.6 27B has the lower benchmark cost at $0.067 vs $0.127. Qwen3.6 27B is faster at 10.65s vs 19.15s, with pass rates of 34.9% vs 45.5%.

Last updated at: 2026-07-25

Rank
#167
Total Output Tokens
26,638
Response Time (avg)
19.15s
Total Cost
$0.127
Rank
#164
Total Output Tokens
16,155
Response Time (avg)
10.65s
Total Cost
$0.067
Recommended model Qwen3.6 27B

It has the best score here (5.5), while costing about 1.9x less than Kimi K2.5.

Detailed comparison

Metric Kimi K2.5 Kimi K2.5 none Release: 2026-01-27 Qwen3.6 27B Qwen3.6 27B none Release: 2026-04-20
Score 5.5 5.5
Rank #167 #164
Reliability 10.0 10.0
Consistency 8.6 7.6
Tests Correct
Attempt pass rate 34.9% 45.5%
Flaky tests 4 6
Total Runs 66 66
Cost per result 1.898 1.220
Total Cost $0.127 $0.067
Input Price $0.571 / 1M $0.290 / 1M
Output Price $2.850 / 1M $2.400 / 1M
Total Input Tokens 89,322 95,796
Output Tokens 26,638 16,155
Reasoning Tokens 0 0
Response Time (avg) 19.15s 10.65s
Response Time (max) 102.83s 156.31s
Response Time (total) 287.30s 234.39s

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#167 MoonshotAI: Kimi K2.5

none
Cost
$0.015
Time
89.1s
Tokens
5,421 tok

#164 Qwen3.6 27B

none
Cost
$0.009
Time
83.0s
Tokens
4,549 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Kimi K2.5 5.5 10.0 33.3% 0 24.56s 7,311 4,708 0
Qwen3.6 27B 5.5 10.0 33.3% 0 4.16s 7,913 539 0

Quick Compare

Switch Comparison Pair