Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: GPT-5.2 Chat vs OpenAI: GPT-5.4

Last updated at: 2026-06-01

Metric GPT-5.2 Chat GPT-5.2 Chat none Release: 2025-12-11 GPT-5.4 GPT-5.4 medium Release: 2026-03-05
Score 7.9 7.9
Rank #32 #29
Reliability 10.0 10.0
Consistency 8.9 8.5
Tests Correct
Attempt pass rate 73.3% 75.0%
Flaky tests 3 4
Total Runs 60 60
Cost per result 2.703 8.765
Total Cost $0.352 $1.140
Input Price $1.750 / 1M $2.500 / 1M
Output Price $14.000 / 1M $15.000 / 1M
Output Tokens 21,144 2,221
Reasoning Tokens 0 68,486
Response Time (avg) 6.82s 22.31s
Response Time (max) 38.52s 100.41s
Response Time (total) 136.34s 446.17s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 8.7 7.9 91.7% 1 3.40s 1,807 0
GPT-5.4 8.3 10.0 75.0% 0 4.11s 240 1,511
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 8.2 6.7 83.3% 1 8.05s 4,131 0
GPT-5.4 8.2 6.7 83.3% 1 54.98s 412 19,995
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 10.0 10.0 100.0% 0 9.12s 1,243 0
GPT-5.4 10.0 10.0 100.0% 0 20.57s 301 3,543
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 10.0 10.0 100.0% 0 3.05s 980 0
GPT-5.4 10.0 10.0 100.0% 0 5.32s 234 804
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 5.3 10.0 33.3% 0 17.78s 7,810 0
GPT-5.4 5.3 7.2 44.4% 1 74.27s 61 34,748
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 4.4 3.0 33.3% 1 3.20s 335 0
GPT-5.4 4.7 3.1 33.3% 1 4.92s 145 321
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 9.8 10.0 100.0% 0 5.51s 1,441 0
GPT-5.4 10.0 10.0 100.0% 0 3.11s 93 897
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 7.7 10.0 66.7% 0 4.10s 1,603 0
GPT-5.4 8.2 7.2 88.9% 1 9.14s 441 3,815
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 10.0 10.0 100.0% 0 4.68s 555 0
GPT-5.4 10.0 10.0 100.0% 0 13.28s 264 1,031
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 Chat 3.0 10.0 0.0% 0 6.89s 1,239 0
GPT-5.4 3.0 10.0 0.0% 0 13.95s 30 1,821

Quick Compare

Switch Comparison Pair