Navigate
Advertise here

Claude Sonnet 5.5 (low) vs Gemini 3.1 Flash Lite Preview (medium)

Gemini 3.1 Flash Lite Preview (medium) leads on average score with 7.0 vs 6.9. Gemini 3.1 Flash Lite Preview (medium) has the lower benchmark cost at $0.158 vs $0.909. Gemini 3.1 Flash Lite Preview (medium) is faster at 6.78s vs 7.46s, with pass rates of 46.4% vs 60.9%.

Last updated at: 2026-10-09

Compared models

Rank
#178
Total Output Tokens
32,443
Response Time (avg)
7.46s
Total Cost
$0.909
Rank
#166
Total Output Tokens
59,534
Response Time (avg)
6.78s
Total Cost
$0.158
Recommended model Gemini 3.1 Flash Lite Preview (medium)

It has the best score here (7.0), while costing about 5.8x less than Claude Sonnet 5.5 (low).

Detailed comparison

Metric Claude Sonnet 5.5 Claude Sonnet 5.5 low Release: 2026-09-29 Gemini 3.1 Flash Lite Preview Gemini 3.1 Flash Lite Preview medium Release: 2026-03-03
Score 6.9 7.0
Rank #178 #166
Reliability 10.0 10.0
Consistency 9.4 9.9
Attempts 69/69 69/69
Tests Correct
Attempt pass rate 46.4% 60.9%
Flaky tests 2 0
Total Runs 69 69
Cost per result 9.082 1.124
Total Cost $0.909 $0.158
Input Price $2.000 / 1M $0.250 / 1M
Output Price $10.000 / 1M $1.500 / 1M
Cache Read Price $0.200 / 1M $0.025 / 1M
Cache Write Price $2.500 / 1M $0.084 / 1M
Total Input Tokens 291,853 272,221
Output Tokens 24,116 12,345
Reasoning Tokens 8,327 47,189
Response Time (avg) 7.46s 6.78s
Response Time (max) 46.61s 54.91s
Response Time (total) 171.47s 155.96s
Parameters ~1T total (~100B active) ~150B total (~10B active)
Availability Closed Closed

Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.

Model generation showcase

Hamster playing table tennis

Prompt: Create a detailed SVG illustration of a hamster playing table tennis.

#178 Claude Sonnet 5.5

low
Cost
$0.020
Time
15.3s
Tokens
2,077 tok

#166 Gemini 3.1 Flash Lite Preview

medium
Cost
$0.003
Time
5.2s
Tokens
1,944 tok

Top Models by Score

Score vs Total Cost

Response Time (avg)

Score vs Response Time (avg)

Total Output Tokens

Score vs Total Output Tokens

Category Breakdown

Coding Score Consistency Attempt pass rate Flaky tests Tests Correct Response Time (avg) Input Tokens Output Tokens Reasoning Tokens
Claude Sonnet 5.5 3.2 9.7 0.0% 0 5.81s 10,608 5,903 0
Gemini 3.1 Flash Lite Preview 5.5 10.0 33.3% 0 4.09s 8,126 461 8,597

Switch Comparison Pair