Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: GPT-5.4 Nano vs xAI: Grok 4.20 Multi-Agent Beta

Last updated at: 2026-03-17

Metric GPT-5.4 Nano GPT-5.4 Nano none Release: 2026-03-17 Grok 4.20 Multi-Agent Beta Grok 4.20 Multi-Agent Beta medium Release: 2026-03-12
Rank #73 #44
Score 4.3 6.2
Consistency 7.3 7.2
Cost per result 0.404 82.962
Total Cost $0.009 $4.978
Tests Correct
Attempt pass rate 29.4% 54.9%
Flaky tests 6 6
Total Runs 51 51
Output Tokens 2,185 298,948
Reasoning Tokens 0 296,529
Response Time (avg) 1.39s 8.64s
Response Time (max) 3.84s 35.28s
Response Time (total) 23.70s 129.64s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 3.5 8.0 16.7% 1 1.18s 800 0
Grok 4.20 Multi-Agent Beta 6.9 5.8 75.0% 2 3.46s 33,706 33,077
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 3.0 10.0 0.0% 0 3.84s 280 0
Grok 4.20 Multi-Agent Beta 3.0 10.0 0.0% 0 0ms 0 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 6.5 10.0 50.0% 0 1.11s 219 0
Grok 4.20 Multi-Agent Beta 10.0 10.0 100.0% 0 5.54s 25,306 25,051
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 2.9 4.4 22.2% 2 926ms 52 0
Grok 4.20 Multi-Agent Beta 2.9 7.2 11.1% 1 24.67s 164,609 163,647
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 3.8 2.5 33.3% 1 1.31s 180 0
Grok 4.20 Multi-Agent Beta 5.8 2.8 66.7% 1 6.40s 15,848 15,746
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 5.0 6.8 33.3% 1 787ms 84 0
Grok 4.20 Multi-Agent Beta 8.3 10.0 50.0% 0 4.63s 25,457 25,322
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 3.7 7.3 22.2% 1 1.29s 348 0
Grok 4.20 Multi-Agent Beta 7.2 5.1 77.8% 2 5.01s 34,022 33,686
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.4 Nano 10.0 10.0 100.0% 0 3.40s 222 0
Grok 4.20 Multi-Agent Beta 3.0 10.0 0.0% 0 0ms 0 0

Quick Compare

Switch Comparison Pair