Navigate
AI BENCHY
Your ad here

AI BENCHY Compare

HY3 Preview vs Grok 4.20 Multi Agent Beta

Last updated at: 2026-04-26

Metric HY3 Preview HY3 Preview high Release: 2026-04-22 Free Available Grok 4.20 Multi Agent Beta Grok 4.20 Multi Agent Beta medium Release: 2026-03-12
Score 8.5 6.4
Rank #11 #67
Reliability N/A N/A
Consistency 8.8 7.4
Tests Correct
Attempt pass rate 81.5% 57.4%
Flaky tests 3 6
Total Runs 50 52
Cost per result 0.000 72.473
Total Cost $0.000 $5.074
Input Price $0.000 / 1M $0.000 / 1M
Output Price $0.000 / 1M $0.000 / 1M
Output Tokens 238,920 299,034
Reasoning Tokens 0 309,670
Response Time (avg) 55.19s 9.80s
Response Time (max) 149.94s 35.28s
Response Time (total) 938.23s 156.75s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 10.0 10.0 100.0% 0 32.69s 26,550 0
Grok 4.20 Multi Agent Beta 6.9 5.8 75.0% 2 3.46s 33,706 33,077
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 10.0 10.0 100.0% 0 99.76s 38,167 0
Grok 4.20 Multi Agent Beta 10.0 10.0 100.0% 0 27.11s 86 13,141
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 10.0 10.0 100.0% 0 113.09s 31,319 0
Grok 4.20 Multi Agent Beta 3.0 10.0 0.0% 0 0ms 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 6.5 10.0 50.0% 0 12.11s 4,323 0
Grok 4.20 Multi Agent Beta 10.0 10.0 100.0% 0 5.54s 25,306 25,051
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 5.3 7.2 44.4% 1 109.04s 87,559 0
Grok 4.20 Multi Agent Beta 2.9 7.2 11.1% 1 24.67s 164,609 163,647
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 10.0 10.0 100.0% 0 24.31s 5,490 0
Grok 4.20 Multi Agent Beta 5.8 2.8 66.7% 1 6.40s 15,848 15,746
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 8.5 6.8 83.3% 1 34.02s 13,331 0
Grok 4.20 Multi Agent Beta 8.3 10.0 50.0% 0 4.63s 25,457 25,322
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 9.0 7.9 88.9% 1 28.07s 21,811 0
Grok 4.20 Multi Agent Beta 7.2 5.1 77.8% 2 5.01s 34,022 33,686
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
HY3 Preview 10.0 10.0 100.0% 0 78.83s 10,370 0
Grok 4.20 Multi Agent Beta 3.0 10.0 0.0% 0 0ms 0 0

Quick Compare

Switch Comparison Pair