Navigate
Advertise here

gpt-oss-120b (medium) vs Inkling Small (medium)

The average score is effectively tied at 5.9 vs 6.0. gpt-oss-120b (medium) has the lower benchmark cost at $0.030 vs $0.113. Inkling Small (medium) is faster at 6.03s vs 31.58s, with pass rates of 47.8% vs 50.7%.

Last updated at: 2026-10-02

Compared models

Rank
#246
Total Output Tokens
118,831
Response Time (avg)
31.58s
Total Cost
$0.030
Rank
#243
Total Output Tokens
56,269
Response Time (avg)
6.03s
Total Cost
$0.113
Recommended model gpt-oss-120b (medium)

It has the best score here (5.9), while costing about 3.9x less than Inkling Small (medium).

Detailed comparison

Metric gpt-oss-120b gpt-oss-120b medium Release: 2025-08-05 Inkling Small Inkling Small medium Release: 2026-08-01 Free Available
Score 5.9 6.0
Rank #246 #243
Reliability 10.0 10.0
Consistency 8.1 9.1
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 47.8% 50.7%
Flaky tests 5 3
Total Runs 69 69
Cost per result 0.335 1.177
Total Cost $0.030 $0.113
Input Price $0.037 / 1M $0.450 / 1M
Output Price $0.170 / 1M $1.200 / 1M
Cache Read Price N/A $0.100 / 1M
Cache Write Price N/A N/A
Total Input Tokens 285,293 100,255
Output Tokens 29,565 6,422
Reasoning Tokens 89,266 49,847
Response Time (avg) 31.58s 6.03s
Response Time (max) 203.90s 17.26s
Response Time (total) 536.78s 138.65s
Parameters 117B total (5.1B active) 276B total (12B active)
Availability Open source Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#246 gpt-oss-120b

medium
Cost
$0.001
Time
26.7s
Tokens
555 tok

#243 Thinking Machines: Inkling Small

medium
Cost
$0.003
Time
13.4s
Tokens
1,753 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
gpt-oss-120b 5.9 7.0 55.6% 1 38.37s 7,782 3,365 11,973
Inkling Small 7.8 10.0 66.7% 0 10.49s 7,374 434 13,783

Quick Compare

Switch Comparison Pair