Navigate
AI BENCHY
Advertise here

Cobuddy (medium) vs Mercury 2

The average score is effectively tied at 4.7 vs 4.7. Cobuddy (medium) has the lower benchmark cost at $0.000 vs $0.030. Mercury 2 is faster at 825ms vs 39.90s, with pass rates of 45.5% vs 24.2%.

Last updated at: 2026-09-03

Rank
#280
Total Output Tokens
118,380
Response Time (avg)
39.90s
Total Cost
$0.000
Rank
#281
Total Output Tokens
9,584
Response Time (avg)
825ms
Total Cost
$0.030
Recommended model Mercury 2

It has the best score here (4.7), while responding about 48.4x faster than Cobuddy (medium).

Detailed comparison

Metric Cobuddy Cobuddy medium Release: 2026-05-06 Mercury 2 Mercury 2 none Release: 2026-02-24
Score 4.7 4.7
Rank #280 #281
Reliability 10.0 10.0
Consistency 7.2 9.2
Attempts 63/66 66/66
Tests Correct
Attempt pass rate 45.5% 24.2%
Flaky tests 6 2
Total Runs 63 66
Cost per result 0.000 0.740
Total Cost $0.000 $0.030
Input Price $0.000 / 1M $0.250 / 1M
Output Price $0.000 / 1M $0.750 / 1M
Total Input Tokens 37,449 89,637
Output Tokens 1,677 9,584
Reasoning Tokens 116,703 0
Response Time (avg) 39.90s 825ms
Response Time (max) 309.02s 4.52s
Response Time (total) 797.98s 18.14s
Parameters ~21B total (~3B active) ~100B
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#280 Cobuddy

medium
No endpoints found for baidu/cobuddy:free.
Cost
$0.000
Time
0.1s
Tokens
0 tok

#281 Mercury 2

none
Cost
$0.002
Time
1.8s
Tokens
1,514 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Cobuddy 3.7 6.7 22.2% 1 79.17s 4,726 358 30,138
Mercury 2 3.4 9.6 0.0% 0 1.03s 7,229 3,088 0

Quick Compare

Switch Comparison Pair