Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: GPT-5.4 vs Qwen: Qwen3.5-Flash

Summary

GPT-5.4 vs Qwen3.5-Flash benchmark comparison: Qwen3.5-Flash leads on average score with 6.1 vs 5.8. Qwen3.5-Flash has the lower benchmark cost at $0.005 vs $0.122. GPT-5.4 is faster at 1.42s vs 3.58s, with pass rates of 36.5% vs 39.7%.

Recommended model: Qwen3.5-Flash - It has the best score here (6.1), while costing about 29.5x less than GPT-5.4.

Last updated at: 2026-06-18

Metric GPT-5.4 GPT-5.4 none Release: 2026-03-05 Qwen3.5-Flash Qwen3.5-Flash none Release: 2026-02-24
Score 5.8 6.1
Rank #112 #97
Reliability 10.0 10.0
Consistency 9.2 9.7
Tests Correct
Attempt pass rate 36.5% 39.7%
Flaky tests 2 1
Total Runs 63 63
Cost per result 1.740 0.075
Total Cost $0.122 $0.005
Input Price $2.500 / 1M $0.065 / 1M
Output Price $15.000 / 1M $0.260 / 1M
Total Input Tokens 34,212 46,439
Output Tokens 2,417 4,276
Reasoning Tokens 0 0
Response Time (avg) 1.42s 3.58s
Response Time (max) 2.95s 27.18s
Response Time (total) 29.87s 75.28s

Generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#112 GPT-5.4

none
Cost
$0.026
Time
18.1s
Tokens
1,792 tok

#97 Qwen3.5-Flash

none
Cost
$0.003
Time
47.4s
Tokens
7,799 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 3.2 8.0 8.3% 1 1.21s 606 406 0
Qwen3.5-Flash 3.5 8.3 8.3% 1 1.32s 696 690 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 5.5 10.0 33.3% 0 1.62s 7,305 516 0
Qwen3.5-Flash 5.5 10.0 33.3% 0 850ms 7,913 519 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 3.0 10.0 0.0% 0 2.89s 11,019 291 0
Qwen3.5-Flash 3.0 10.0 0.0% 0 6.22s 18,879 1,794 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 10.0 10.0 100.0% 0 1.04s 7,140 222 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 1.57s 7,794 243 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 5.3 7.2 44.4% 1 1.07s 723 50 0
Qwen3.5-Flash 7.7 10.0 66.7% 0 905ms 789 15 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 4.4 9.9 0.0% 0 1.78s 477 184 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 803ms 522 100 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 6.5 10.0 50.0% 0 1.07s 660 81 0
Qwen3.5-Flash 6.3 10.0 50.0% 0 8.81s 711 63 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 5.6 9.8 33.3% 0 1.44s 642 381 0
Qwen3.5-Flash 3.1 10.0 0.0% 0 10.89s 714 579 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 10.0 10.0 100.0% 0 2.75s 5,445 246 0
Qwen3.5-Flash 10.0 10.0 100.0% 0 3.67s 8,211 264 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.4 3.0 10.0 0.0% 0 990ms 195 40 0
Qwen3.5-Flash 3.0 10.0 0.0% 0 588ms 210 9 0

Quick Compare

Switch Comparison Pair