Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

Inception: Mercury 2 vs Laguna M.1

Last updated at: 2026-04-29

Metric Mercury 2 Mercury 2 medium Release: 2026-02-24 Laguna M.1 Laguna M.1 medium Release: 2026-04-28 Free Available
Score 6.5 6.3
Rank #71 #74
Reliability N/A 10.0
Consistency 8.6 8.6
Tests Correct
Attempt pass rate 53.7% 53.7%
Flaky tests 3 3
Total Runs 54 54
Cost per result 0.580 0.000
Total Cost $0.047 $0.000
Input Price $0.250 / 1M $0.000 / 1M
Output Price $0.750 / 1M $0.000 / 1M
Output Tokens 3,972 63,822
Reasoning Tokens 48,333 0
Response Time (avg) 2.21s 13.90s
Response Time (max) 14.63s 53.14s
Response Time (total) 37.51s 250.28s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 6.9 9.9 50.0% 0 1.12s 2,546 2,609
Laguna M.1 6.6 10.0 50.0% 0 9.15s 7,839 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 10.0 10.0 100.0% 0 1.53s 249 2,213
Laguna M.1 4.3 1.1 66.7% 1 35.61s 14,327 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 10.0 10.0 100.0% 0 3.28s 268 4,887
Laguna M.1 3.0 10.0 0.0% 0 53.14s 12,272 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 7.3 5.9 83.3% 1 1.11s 183 1,656
Laguna M.1 10.0 10.0 100.0% 0 4.93s 2,296 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 2.9 7.2 11.1% 1 6.48s 41 30,754
Laguna M.1 5.3 7.2 44.4% 1 24.14s 19,020 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 4.8 10.0 0.0% 0 821ms 137 542
Laguna M.1 4.1 10.0 0.0% 0 6.86s 1,294 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 10.0 10.0 100.0% 0 1.07s 14 958
Laguna M.1 10.0 10.0 100.0% 0 4.30s 1,626 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 3.9 7.5 22.2% 1 934ms 354 2,758
Laguna M.1 3.6 7.2 22.2% 1 6.97s 3,978 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Mercury 2 10.0 10.0 100.0% 0 1.89s 180 1,956
Laguna M.1 10.0 10.0 100.0% 0 6.31s 1,170 0

Quick Compare

Switch Comparison Pair