Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

inclusionAI: Ring-2.6-1T vs xAI: Grok 4.3

Last updated at: 2026-05-22

Metric Ring-2.6-1T Ring-2.6-1T medium Release: 2026-05-10 Grok 4.3 Grok 4.3 medium Release: 2026-05-01
Score 7.2 7.8
Rank #61 #31
Reliability 9.9 10.0
Consistency 8.7 8.4
Tests Correct
Attempt pass rate 66.7% 75.0%
Flaky tests 3 4
Total Runs 60 60
Cost per result 0.000 4.562
Total Cost $0.000 $0.593
Input Price $0.075 / 1M $1.250 / 1M
Output Price $0.625 / 1M $2.500 / 1M
Output Tokens 21,752 1,485
Reasoning Tokens 42,754 214,928
Response Time (avg) 61.29s 49.23s
Response Time (max) 304.19s 216.69s
Response Time (total) 1164.50s 984.54s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 10.0 10.0 100.0% 0 42.21s 3,833 4,891
Grok 4.3 10.0 10.0 100.0% 0 8.83s 88 8,207
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 10.0 10.0 100.0% 0 59.65s 1,369 3,985
Grok 4.3 7.4 6.5 66.7% 1 55.26s 532 24,554
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 10.0 10.0 100.0% 0 304.19s 324 6,088
Grok 4.3 10.0 10.0 100.0% 0 63.99s 234 15,301
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 6.5 10.0 50.0% 0 37.36s 840 1,937
Grok 4.3 10.0 10.0 100.0% 0 18.97s 180 9,546
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 3.5 4.4 33.3% 2 64.92s 9,744 15,013
Grok 4.3 5.3 7.2 44.4% 1 181.74s 14 111,300
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 4.1 10.0 0.0% 0 58.26s 150 583
Grok 4.3 5.4 2.5 66.7% 1 24.70s 70 5,020
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 9.8 10.0 100.0% 0 11.78s 266 1,831
Grok 4.3 9.8 10.0 100.0% 0 18.58s 57 8,713
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 5.9 7.2 55.6% 1 20.73s 697 2,479
Grok 4.3 5.9 7.2 55.6% 1 22.53s 128 14,686
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 10.0 10.0 100.0% 0 104.44s 234 1,531
Grok 4.3 10.0 10.0 100.0% 0 17.66s 168 4,615
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
Ring-2.6-1T 3.0 10.0 0.0% 0 113.91s 4,295 4,416
Grok 4.3 3.0 10.0 0.0% 0 44.47s 14 12,986

Quick Compare

Switch Comparison Pair