Navigate
AI BENCHY
Compare Charts
❤️ Made by XCS
Your ad here

AI BENCHY Compare

Inception: Mercury 2 vs xAI: Grok 4.1 Fast

Compare:

Last updated at: 2026-03-05

Metric Inception: Mercury 2 medium Release: 2026-02-24 xAI: Grok 4.1 Fast none Release: 2025-11-19
Rank #35 #53
Avg Score 5.4 2.9
Tests Correct
Consistency 8.3 8.9
Cost per result 0.622 0.239
Total Cost $0.044 $0.008
Attempt pass rate 57.8% 26.7%
Flaky tests 3 2
common.totalAttempts 45 (15 x 3) 45 (15 x 3)
Output Tokens 3,571 1,036
Reasoning Tokens 45,379 0
Response Time (avg) 2.47s 2.01s
Response Time (max) 14.63s 5.51s
Response Time (total) 34.56s 16.06s

Top Models by Score

Response Time (avg)

Score vs Total Cost

Avg Score vs Response Time (avg)

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 7.3 9.8 66.7% 0 1.30s 2,531 2,410
xAI: Grok 4.1 Fast 1.3 10.0 0.0% 0 1.73s 229 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 10.0 10.0 100.0% 0 3.28s 268 4,887
xAI: Grok 4.1 Fast 10.0 10.0 0.0% 0 3.33s 105 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 5.5 5.9 83.3% 1 1.11s 183 1,656
xAI: Grok 4.1 Fast 9.9 10.0 100.0% 0 943ms 180 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 10.0 7.2 11.1% 1 6.48s 41 30,754
xAI: Grok 4.1 Fast 4.0 7.2 55.6% 1 1.06s 15 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 10.0 10.0 100.0% 0 1.07s 14 958
xAI: Grok 4.1 Fast 10.0 10.0 0.0% 0 923ms 56 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 1.7 7.5 22.2% 1 934ms 354 2,758
xAI: Grok 4.1 Fast 1.3 10.0 0.0% 0 1.28s 243 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Inception: Mercury 2 10.0 10.0 100.0% 0 1.89s 180 1,956
xAI: Grok 4.1 Fast 10.0 1.6 33.3% 1 5.51s 208 0

Quick Compare

Switch Comparison Pair