Navigate
AI BENCHY
Advertise here

AI BENCHY Compare

OpenAI: GPT-5.2 vs Qwen: Qwen3.6 Flash

Last updated at: 2026-05-19

Metric GPT-5.2 GPT-5.2 medium Release: 2025-12-11 Qwen3.6 Flash Qwen3.6 Flash medium Release: 2026-04-20
Score 7.2 7.5
Rank #65 #55
Reliability 10.0 10.0
Consistency 8.2 7.9
Tests Correct
Attempt pass rate 68.4% 71.9%
Flaky tests 4 5
Total Runs 57 57
Cost per result 3.609 2.768
Total Cost $0.397 $0.305
Input Price $1.750 / 1M $0.188 / 1M
Output Price $14.000 / 1M $1.125 / 1M
Output Tokens 2,731 2,830
Reasoning Tokens 22,200 194,258
Response Time (avg) 15.22s 15.85s
Response Time (max) 77.80s 122.87s
Response Time (total) 182.59s 301.13s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 6.5 8.0 58.3% 1 7.81s 567 2,002
Qwen3.6 Flash 10.0 10.0 100.0% 0 6.10s 624 14,024
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 10.0 10.0 100.0% 0 15.12s 467 2,166
Qwen3.6 Flash 6.7 3.5 66.7% 1 25.84s 435 17,044
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 10.0 10.0 100.0% 0 14.06s 291 1,757
Qwen3.6 Flash 10.0 10.0 100.0% 0 20.28s 483 13,839
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 10.0 10.0 100.0% 0 3.15s 234 420
Qwen3.6 Flash 10.0 10.0 100.0% 0 9.65s 270 13,155
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 5.9 7.2 55.6% 1 77.80s 42 10,342
Qwen3.6 Flash 3.5 4.4 33.3% 2 14.65s 60 24,409
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 3.7 9.7 0.0% 0 4.32s 162 269
Qwen3.6 Flash 4.8 9.9 0.0% 0 9.88s 140 5,445
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 9.9 10.0 100.0% 0 3.12s 94 614
Qwen3.6 Flash 10.0 10.0 100.0% 0 6.05s 102 7,423
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 7.6 7.3 77.8% 1 5.47s 609 938
Qwen3.6 Flash 6.1 4.7 66.7% 2 6.17s 355 10,683
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 4.7 1.6 66.7% 1 10.30s 239 469
Qwen3.6 Flash 10.0 10.0 100.0% 0 4.00s 335 1,188
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
GPT-5.2 3.0 10.0 0.0% 0 28.18s 26 3,223
Qwen3.6 Flash 3.0 10.0 0.0% 0 122.87s 26 87,048

Quick Compare

Switch Comparison Pair