Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

Cobuddy vs MoonshotAI: Kimi K2.5

Summary

Cobuddy vs Kimi K2.5 benchmark comparison: Kimi K2.5 leads on average score with 5.5 vs 4.9. Cobuddy has the lower benchmark cost at $0.000 vs $0.027. Kimi K2.5 is faster at 13.18s vs 39.90s, with pass rates of 47.6% vs 34.9%.

Recommended model: Kimi K2.5 - It has the best score here (5.5), while responding about 3.0x faster than Cobuddy.

Last updated at: 2026-07-02

Metric Cobuddy Cobuddy medium Release: 2026-05-06 Kimi K2.5 Kimi K2.5 none Release: 2026-01-27
Score 4.9 5.5
Rank #145 #122
Reliability 10.0 10.0
Consistency 7.5 8.9
Tests Correct
Attempt pass rate 47.6% 34.9%
Flaky tests 6 3
Total Runs 63 63
Cost per result 0.000 0.442
Total Cost $0.000 $0.027
Input Price $0.000 / 1M $0.375 / 1M
Output Price $0.000 / 1M $2.025 / 1M
Total Input Tokens 37,449 36,034
Output Tokens 1,677 6,657
Reasoning Tokens 116,703 0
Response Time (avg) 39.90s 13.18s
Response Time (max) 309.02s 42.13s
Response Time (total) 797.98s 184.47s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#145 Cobuddy

medium
No endpoints found for baidu/cobuddy:free.
Cost
$0.000
Time
0.1s
Tokens
0 tok

#122 MoonshotAI: Kimi K2.5

none
Cost
$0.015
Time
89.1s
Tokens
5,421 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 8.7 7.9 91.7% 1 10.00s 453 98 4,666
Kimi K2.5 3.6 8.4 8.3% 1 6.24s 652 373 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 3.7 6.7 22.2% 1 79.17s 4,726 358 30,138
Kimi K2.5 5.5 10.0 33.3% 0 24.56s 7,311 4,708 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 3.0 10.0 0.0% 0 47.38s 18,324 465 7,265
Kimi K2.5 2.8 2.1 33.3% 1 19.16s 12,264 748 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 6.3 5.8 66.7% 1 17.36s 8,181 275 5,591
Kimi K2.5 7.3 5.8 83.3% 1 42.13s 7,180 187 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 2.9 4.4 22.2% 2 128.15s 540 10 49,454
Kimi K2.5 5.3 10.0 33.3% 0 4.38s 753 29 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 4.2 9.9 0.0% 0 23.23s 498 76 3,782
Kimi K2.5 10.0 10.0 100.0% 0 4.00s 483 76 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 9.8 10.0 100.0% 0 11.60s 508 64 2,842
Kimi K2.5 6.5 10.0 50.0% 0 2.67s 677 60 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 3.6 7.2 22.2% 1 12.83s 561 189 5,808
Kimi K2.5 3.0 10.0 0.0% 0 4.04s 667 236 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 10.0 10.0 100.0% 0 11.19s 3,505 133 294
Kimi K2.5 10.0 10.0 100.0% 0 13.99s 5,835 220 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 3.0 10.0 0.0% 0 36.98s 153 9 6,863
Kimi K2.5 3.0 10.0 0.0% 0 3.90s 212 20 0

Quick Compare

Switch Comparison Pair