- Rank
- #293
- Total Output Tokens
- 40,302
- Response Time (avg)
- 3.34s
- Total Cost
- $0.020
Mercury 2.5 (low) vs Nemotron 3.5 Lightning (high)
Nemotron 3.5 Lightning (high) leads on average score with 5.2 vs 5.1. Mercury 2.5 (low) has the lower benchmark cost at $0.020 vs $0.111. Mercury 2.5 (low) is faster at 3.34s vs 77.08s, with pass rates of 37.7% vs 47.8%.
Compared models
- Rank
- #287
- Total Output Tokens
- 659,693
- Response Time (avg)
- 77.08s
- Total Cost
- $0.111
Recommended model
Mercury 2.5 (low)
Its score stays close to the best score here (5.1 vs 5.2), while costing about 5.8x less than Nemotron 3.5 Lightning (high).
Detailed comparison
| Metric | Mercury 2.5 Mercury 2.5 low | Nemotron 3.5 Lightning Nemotron 3.5 Lightning high Free Available |
|---|---|---|
| Score | 5.1 | 5.2 |
| Rank | #293 | #287 |
| Reliability | 9.8 | 9.6 |
| Consistency | 8.2 | 5.7 |
| Attempts | 69/69 | 69/69 |
| Tests Correct | ||
| Attempt pass rate | 37.7% | 47.8% |
| Flaky tests | 5 | 12 |
| Total Runs | 69 | 69 |
| Cost per result | 0.319 | 0.000 |
| Total Cost | $0.020 | $0.111 |
| Input Price | $0.040 / 1M | $0.060 / 1M |
| Output Price | $0.150 / 1M | $0.160 / 1M |
| Total Input Tokens | 326,083 | 170,517 |
| Output Tokens | 7,224 | 149,968 |
| Reasoning Tokens | 33,078 | 509,725 |
| Response Time (avg) | 3.34s | 77.08s |
| Response Time (max) | 47.06s | 490.99s |
| Response Time (total) | 76.80s | 1772.78s |
| Parameters | ~100B | 30B total (3B active) |
| Availability | Closed | Weights available |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#293 Mercury 2.5
low- Cost
- $0.001
- Time
- 2.4s
- Tokens
- 1,236 tok
#287 Nemotron 3.5 Lightning
high- Cost
- $0.000
- Time
- 92.8s
- Tokens
- 11,343 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Agentic | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 5.0 | 10.0 | 0.0% | 0 | 47.06s | 188,063 | 1,141 | 5,442 | |
| Nemotron 3.5 Lightning | 3.5 | 8.3 | 0.0% | 0 | 490.99s | 52,501 | 1,624 | 3,472 |
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 6.9 | 7.9 | 66.7% | 1 | 799ms | 772 | 209 | 2,969 | |
| Nemotron 3.5 Lightning | 5.5 | 3.7 | 66.7% | 3 | 10.29s | 696 | 2,432 | 15,948 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 5.5 | 10.0 | 33.3% | 0 | 972ms | 7,909 | 521 | 2,269 | |
| Nemotron 3.5 Lightning | 4.4 | 5.1 | 33.3% | 2 | 168.41s | 7,623 | 68,533 | 242,319 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 3.0 | 10.0 | 0.0% | 0 | 5.16s | 107,683 | 3,750 | 11,351 | |
| Nemotron 3.5 Lightning | 6.4 | 10.0 | 50.0% | 0 | 160.45s | 90,257 | 17,550 | 119,725 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 6.5 | 10.0 | 50.0% | 0 | 1.06s | 8,322 | 739 | 1,499 | |
| Nemotron 3.5 Lightning | 7.3 | 5.9 | 83.3% | 1 | 13.99s | 7,944 | 3,210 | 16,835 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 3.6 | 7.2 | 22.2% | 1 | 846ms | 851 | 57 | 2,265 | |
| Nemotron 3.5 Lightning | 2.9 | 4.4 | 22.2% | 2 | 88.61s | 798 | 45,616 | 72,496 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 4.4 | 9.9 | 0.0% | 0 | 819ms | 525 | 99 | 748 | |
| Nemotron 3.5 Lightning | 3.7 | 9.5 | 0.0% | 0 | 7.58s | 516 | 4,331 | 4,408 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 2.9 | 6.0 | 16.7% | 1 | 972ms | 518 | 48 | 1,026 | |
| Nemotron 3.5 Lightning | 8.5 | 6.8 | 83.3% | 1 | 4.98s | 723 | 502 | 3,779 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 5.9 | 4.5 | 66.7% | 2 | 792ms | 768 | 249 | 2,214 | |
| Nemotron 3.5 Lightning | 3.6 | 1.8 | 44.4% | 3 | 6.09s | 726 | 1,236 | 9,851 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 9.8 | 10.0 | 100.0% | 0 | 2.72s | 10,440 | 388 | 2,537 | |
| Nemotron 3.5 Lightning | 10.0 | 10.0 | 100.0% | 0 | 6.92s | 8,526 | 3,745 | 3,781 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Mercury 2.5 | 3.0 | 10.0 | 0.0% | 0 | 782ms | 232 | 23 | 758 | |
| Nemotron 3.5 Lightning | 3.0 | 10.0 | 0.0% | 0 | 77.98s | 207 | 1,189 | 17,111 |
Quick Compare
Switch Comparison Pair
Mercury 2.5lowvsKAT Coder AIR V2.5mediumMercury 2.5lowvsMiMo-V2.5-PrononeLing 3.0 FlashnonevsNemotron 3.5 LightninghighFree AvailableMercury 2.5lowvsMistral Small 4noneMercury 2.5lowvsMistral Small 4mediumNemotron 3.5 LightninghighFree AvailablevsInkling SmalllowFree AvailableMistral Small 4mediumvsNemotron 3.5 LightninghighFree AvailableMercury 2.5lowvsLaguna S 2.1mediumFree AvailableMistral Small 4nonevsNemotron 3.5 LightninghighFree AvailableMercury 2.5lowvsKAT Coder AIR V2.5highNemotron 3.5 LightninghighFree AvailablevsQwen3.6 35B A3BnoneNemotron 3.5 LightninghighFree AvailablevsLaguna XS 2.1noneFree Available