Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

GPT-5.6 Terra (medium) vs Qwen3.7 Flash (high)

GPT-5.6 Terra (medium) leads on average score with 7.8 vs 7.8. Qwen3.7 Flash (high) has the lower benchmark cost at $0.052 vs $0.523. GPT-5.6 Terra (medium) is faster at 6.85s vs 41.37s, with pass rates of 69.7% vs 71.2%.

Last updated at: 2026-09-10

Compared models

Rank
#86
Total Output Tokens
30,349
Response Time (avg)
6.85s
Total Cost
$0.523
Rank
#91
Total Output Tokens
370,250
Response Time (avg)
41.37s
Total Cost
$0.052
Recommended model Qwen3.7 Flash (high)

Its score stays close to the best score here (7.8 vs 7.8), while costing about 10.1x less than GPT-5.6 Terra (medium).

Detailed comparison

Metric GPT-5.6 Terra GPT-5.6 Terra medium Release: 2026-07-09 Qwen3.7 Flash Qwen3.7 Flash high Release: 2026-07-28
Score 7.8 7.8
Rank #86 #91
Reliability 10.0 10.0
Consistency 9.3 7.8
Attempts 66/66 66/66
Tests Correct
Attempt pass rate 69.7% 71.2%
Flaky tests 2 6
Total Runs 66 66
Cost per result 4.603 0.398
Total Cost $0.523 $0.052
Input Price $2.000 / 1M $0.030 / 1M
Output Price $12.000 / 1M $0.130 / 1M
Total Input Tokens 79,184 116,331
Output Tokens 4,878 12,330
Reasoning Tokens 25,471 357,920
Response Time (avg) 6.85s 41.37s
Response Time (max) 41.68s 431.47s
Response Time (total) 150.65s 910.11s
Parameters ~1T total (~60B active) ~35B total (~3B active)
Availability Closed Closed

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#86 GPT-5.6 Terra

medium
Cost
$0.035
Time
16.3s
Tokens
2,418 tok

#91 Qwen3.7 Flash

high
Reached the allocated time limit (600 seconds) without receiving showcase output.
Cost
$0.000
Time
600.0s
Tokens
0 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-5.6 Terra 6.1 7.2 55.6% 1 7.19s 7,302 427 4,527
Qwen3.7 Flash 8.2 9.7 66.7% 0 58.84s 7,893 503 62,350

Quick Compare

Switch Comparison Pair