Navigate
AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

Compared models

Nemotron 3 Super (medium) vs Qwen3.5-122B-A10B (medium) vs Elephant Alpha (medium) vs gpt-oss-120b (medium) benchmark comparison: Qwen3.5-122B-A10B (medium) leads on Score with 7.1. Qwen3.5-122B-A10B (medium) leads on Reliability with 10.0. Elephant Alpha (medium) has the lowest Total Cost at $0.000. Elephant Alpha (medium) is fastest at 1.27s.

Last updated at: 2026-09-04

Rank
#217
Total Output Tokens
115,574
Response Time (avg)
52.31s
Total Cost
$0.050
Rank
#127
Total Output Tokens
487,519
Response Time (avg)
65.67s
Total Cost
$1.207
Rank
#295
Total Output Tokens
2,596
Response Time (avg)
1.27s
Total Cost
$0.000
Rank
#189
Total Output Tokens
98,282
Response Time (avg)
20.81s
Total Cost
$0.020
Recommended model gpt-oss-120b (medium)

It offers the best overall trade-off: a competitive score (6.1), lower cost than the other models in this comparison, and balanced response time.

Detailed comparison

Metric Nemotron 3 Super Nemotron 3 Super medium Release: 2026-03-11 Free Available Qwen3.5-122B-A10B Qwen3.5-122B-A10B medium Release: 2026-02-24 Elephant Alpha Elephant Alpha medium Release: 2026-04-14 gpt-oss-120b gpt-oss-120b medium Release: 2025-08-05
Score 5.7 7.1 4.3 6.1
Rank #217 #127 #295 #189
Reliability 9.1 10.0 N/A 10.0
Consistency 8.9 8.5 9.2 8.0
Attempts 66/66 66/66 63/66 66/66
Tests Correct
Attempt pass rate 40.9% 71.2% 28.8% 50.0%
Flaky tests 3 4 1 5
Total Runs 66 66 63 66
Cost per result 0.004 8.446 0.000 0.224
Total Cost $0.050 $1.207 $0.000 $0.020
Input Price $0.085 / 1M $0.290 / 1M $0.000 / 1M $0.037 / 1M
Output Price $0.400 / 1M $2.400 / 1M $0.000 / 1M $0.170 / 1M
Total Input Tokens 81,438 124,780 33,744 108,767
Output Tokens 17,674 47,694 2,596 29,378
Reasoning Tokens 97,900 439,825 0 68,904
Response Time (avg) 52.31s 65.67s 1.27s 20.81s
Response Time (max) 431.98s 519.30s 3.70s 68.16s
Response Time (total) 1046.21s 1444.68s 22.82s 332.88s
Parameters 120B total (12B active) 122B total (10B active) 104B total (7.4B active) 117B total (5.1B active)
Availability Weights available Open source Open source Open source

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#217 Nemotron 3 Super

medium
Cost
$0.000
Time
272.6s
Tokens
5,296 tok

#127 Qwen3.5-122B-A10B

medium
Cost
$0.019
Time
48.7s
Tokens
6,034 tok

#295 Elephant Alpha

medium
Elephant Alpha was a stealth model revealed on April 21st as Ling-2.6-flash. Find it here: https://openrouter.ai/inclusionai/ling-2.6-flash:free
Cost
$0.000
Time
0.1s
Tokens
0 tok

#189 gpt-oss-120b

medium
Cost
$0.001
Time
26.7s
Tokens
555 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Nemotron 3 Super 3.1 10.0 0.0% 0 147.32s 2,275 797 4,424
Qwen3.5-122B-A10B 6.0 7.2 55.6% 1 114.48s 7,630 8,057 82,578
Elephant Alpha 3.7 7.8 11.1% 1 1.30s 813 365 0
gpt-oss-120b 5.9 7.0 55.6% 1 38.37s 7,782 3,365 11,973

Quick Compare

Switch Comparison Pair