- Rank
- #47
- Total Output Tokens
- 42,146
- Response Time (avg)
- 15.01s
- Total Cost
- $3.004
Claude Fable 5 (medium) vs DeepSeek V4.1 Flash (low)
Claude Fable 5 (medium) leads on average score with 8.6 vs 8.5. DeepSeek V4.1 Flash (low) has the lower benchmark cost at $0.224 vs $3.004. Claude Fable 5 (medium) is faster at 15.01s vs 26.77s, with pass rates of 78.8% vs 74.2%.
Compared models
- Rank
- #51
- Total Output Tokens
- 347,963
- Response Time (avg)
- 26.77s
- Total Cost
- $0.224
Recommended model
DeepSeek V4.1 Flash (low)
Its score stays close to the best score here (8.5 vs 8.6), while costing about 13.4x less than Claude Fable 5 (medium).
Detailed comparison
| Metric | Claude Fable 5 Claude Fable 5 medium | DeepSeek V4.1 Flash DeepSeek V4.1 Flash low |
|---|---|---|
| Score | 8.6 | 8.5 |
| Rank | #47 | #51 |
| Reliability | 10.0 | 9.5 |
| Consistency | 9.6 | 8.9 |
| Attempts | 66/66 | 66/66 |
| Tests Correct | ||
| Attempt pass rate | 78.8% | 74.2% |
| Flaky tests | 1 | 3 |
| Total Runs | 66 | 66 |
| Cost per result | 17.669 | 1.492 |
| Total Cost | $3.004 | $0.224 |
| Input Price | $10.000 / 1M | $0.150 / 1M |
| Output Price | $50.000 / 1M | $0.600 / 1M |
| Total Input Tokens | 89,640 | 99,715 |
| Output Tokens | 33,092 | 6,381 |
| Reasoning Tokens | 9,054 | 341,582 |
| Response Time (avg) | 15.01s | 26.77s |
| Response Time (max) | 80.80s | 149.84s |
| Response Time (total) | 330.15s | 589.01s |
| Parameters | ~9.5T total (~878B active) | 748B total (16B active) |
| Availability | Closed | Open source |
Model generation showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#47 Claude Fable 5
medium- Cost
- $0.606
- Time
- 156.7s
- Tokens
- 12,264 tok
#51 DeepSeek V4.1 Flash
low- Cost
- $0.017
- Time
- 43.6s
- Tokens
- 14,069 tok
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Category Breakdown
| Anti-AI Tricks | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 6.20s | 834 | 530 | 402 | |
| DeepSeek V4.1 Flash | 8.3 | 10.0 | 75.0% | 0 | 3.32s | 852 | 171 | 4,328 |
| Coding | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 15.59s | 10,590 | 7,383 | 1,318 | |
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 40.66s | 7,509 | 378 | 81,978 |
| Combined | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 6.5 | 10.0 | 50.0% | 0 | 27.47s | 52,197 | 2,373 | 1,599 | |
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 24.73s | 77,368 | 4,906 | 24,016 |
| Data parsing and extraction | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 7.18s | 10,503 | 521 | 363 | |
| DeepSeek V4.1 Flash | 6.5 | 10.0 | 50.0% | 0 | 3.80s | 2,397 | 120 | 1,149 |
| Domain specific | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 5.3 | 7.2 | 44.4% | 1 | 37.32s | 972 | 16,947 | 3,786 | |
| DeepSeek V4.1 Flash | 3.5 | 4.4 | 33.3% | 2 | 95.58s | 909 | 27 | 184,701 |
| General Intelligence | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 7.42s | 708 | 366 | 144 | |
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 4.34s | 549 | 117 | 1,319 |
| Instructions following | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 5.90s | 909 | 139 | 202 | |
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 8.38s | 783 | 63 | 1,546 |
| Puzzle Solving | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 7.7 | 10.0 | 66.7% | 0 | 5.18s | 894 | 402 | 324 | |
| DeepSeek V4.1 Flash | 8.2 | 7.7 | 77.8% | 1 | 9.39s | 828 | 273 | 14,502 |
| Tool Calling | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 10.0 | 10.0 | 100.0% | 0 | 16.96s | 11,775 | 729 | 344 | |
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 13.83s | 8,259 | 313 | 284 |
| Trivia | Score | Consistency | Attempt pass rate | Flaky tests | Tests Correct | Response Time (avg) | Input Tokens | Output Tokens | Reasoning Tokens |
|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5 | 3.0 | 10.0 | 0.0% | 0 | 25.64s | 258 | 3,702 | 572 | |
| DeepSeek V4.1 Flash | 3.0 | 10.0 | 0.0% | 0 | 46.86s | 261 | 13 | 27,759 |
Quick Compare
Switch Comparison Pair
DeepSeek V4.1 FlashlowvsGrok 4.5mediumClaude Fable 5mediumvsMuse Spark 1.2highDeepSeek V4.1 FlashlowvsGrok 4.6mediumDeepSeek V4.1 FlashlowvsGPT-5.4mediumClaude Fable 5mediumvsQwen3.8 MaxlowDeepSeek V4.1 FlashlowvsGPT-5.2mediumDeepSeek V4.1 FlashlowvsQwen3.8 27BhighClaude Fable 5mediumvsGLM 5.3 FlashmaxSeed 2.1 TurbomediumvsDeepSeek V4.1 FlashlowClaude Fable 5mediumvsGemini 3.6 FlashlowDeepSeek V4.1 FlashlowvsQwen3.6 Max PreviewmediumDeepSeek V4.1 FlashlowvsInkling Smallhigh