Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

AI BENCHY Compare

OpenAI: gpt-oss-120b vs Qwen: Qwen3.5 Plus 2026-04-20

Last updated at: 2026-05-26

Metric gpt-oss-120b gpt-oss-120b none Release: 2025-08-05 Free Available Qwen3.5 Plus 2026-04-20 Qwen3.5 Plus 2026-04-20 none Release: 2026-04-20
Score 5.4 5.8
Rank #119 #103
Reliability 10.0 9.9
Consistency 9.1 8.5
Tests Correct
Attempt pass rate 38.6% 43.3%
Flaky tests 2 4
Total Runs 57 60
Cost per result 0.168 0.582
Total Cost $0.011 $0.041
Input Price $0.000 / 1M $0.300 / 1M
Output Price $0.000 / 1M $1.800 / 1M
Output Tokens 51,664 11,139
Reasoning Tokens 0 0
Response Time (avg) 21.61s 4.57s
Response Time (max) 113.71s 33.34s
Response Time (total) 345.79s 91.37s

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Anti-AI Tricks Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.5 10.0 50.0% 0 32.84s 8,676 0
Qwen3.5 Plus 2026-04-20 4.8 10.0 25.0% 0 1.88s 557 0
Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.3 1.1 66.7% 1 9.57s 3,232 0
Qwen3.5 Plus 2026-04-20 4.4 6.7 16.7% 1 2.08s 474 0
Combined Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 0ms 0 0
Qwen3.5 Plus 2026-04-20 2.8 1.6 33.3% 1 13.32s 2,275 0
Data parsing and extraction Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.5 10.0 50.0% 0 7.12s 598 0
Qwen3.5 Plus 2026-04-20 10.0 10.0 100.0% 0 2.82s 243 0
Domain specific Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 34.98s 29,483 0
Qwen3.5 Plus 2026-04-20 5.3 10.0 33.3% 0 4.43s 18 0
General Intelligence Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 4.8 10.0 0.0% 0 10.79s 615 0
Qwen3.5 Plus 2026-04-20 4.8 10.0 0.0% 0 1.41s 119 0
Instructions following Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 9.8 10.0 100.0% 0 5.06s 1,940 0
Qwen3.5 Plus 2026-04-20 6.2 5.8 66.7% 1 1.17s 68 0
Puzzle Solving Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 6.0 7.2 55.6% 1 8.21s 3,982 0
Qwen3.5 Plus 2026-04-20 6.7 7.9 55.6% 1 1.97s 583 0
Tool Calling Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 0ms 0 0
Qwen3.5 Plus 2026-04-20 10.0 10.0 100.0% 0 4.42s 297 0
Trivia Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Output Tokens Reasoning Tokens
gpt-oss-120b 3.0 10.0 0.0% 0 47.29s 3,138 0
Qwen3.5 Plus 2026-04-20 3.0 10.0 0.0% 0 33.34s 6,505 0

Quick Compare

Switch Comparison Pair