Navigate
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

GPT-4o-mini vs Laguna XS 2.1

Laguna XS 2.1 leads on average score with 5.3 vs 5.0. Laguna XS 2.1 has the lower benchmark cost at $0.020 vs $0.024. GPT-4o-mini is faster at 4.16s vs 4.56s, with pass rates of 21.7% vs 29.0%.

Last updated at: 2026-10-02

Compared models

Rank
#304
Total Output Tokens
3,627
Response Time (avg)
4.16s
Total Cost
$0.024
Rank
#286
Total Output Tokens
23,173
Response Time (avg)
4.56s
Total Cost
$0.020
Recommended model Laguna XS 2.1

It has the strongest score in this comparison (5.3) and the best overall balance of cost and response time across all 2 models.

Detailed comparison

Metric GPT-4o-mini GPT-4o-mini none Release: 2024-07-18 Laguna XS 2.1 Laguna XS 2.1 none Release: 2026-07-02 Free Available
Score 5.0 5.3
Rank #304 #286
Reliability 10.0 10.0
Consistency 9.9 9.1
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 21.7% 29.0%
Flaky tests 0 3
Total Runs 69 69
Cost per result 0.480 0.380
Total Cost $0.024 $0.020
Input Price $0.150 / 1M $0.060 / 1M
Output Price $0.600 / 1M $0.120 / 1M
Cache Read Price $0.075 / 1M $0.030 / 1M
Cache Write Price N/A N/A
Total Input Tokens 145,180 270,072
Output Tokens 3,627 23,173
Reasoning Tokens 0 0
Response Time (avg) 4.16s 4.56s
Response Time (max) 39.99s 70.79s
Response Time (total) 70.70s 104.80s
Parameters ~8B 33B total (3B active)
Availability Closed Weights available

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#304 GPT-4o-mini

none
Cost
$0.001
Time
6.6s
Tokens
742 tok

#286 Laguna XS 2.1

none
Cost
$0.001
Time
27.6s
Tokens
4,344 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
GPT-4o-mini 3.2 9.6 0.0% 0 1.63s 7,314 367 0
Laguna XS 2.1 4.3 7.8 22.2% 1 623ms 7,995 562 0

Quick Compare

Switch Comparison Pair