- Rang
- #52
- Totaal aantal uitvoer-tokens
- 452,940
- Responstijd (gem.)
- 32.78s
- Totale kosten
- $0.287
DeepSeek V4.1 Flash (high) vs Qwen3.8 Max (low)
Qwen3.8 Max (low) leidt in gemiddelde score met 8.6 vs 8.5. DeepSeek V4.1 Flash (high) heeft lagere benchmarkkosten met $0.287 vs $0.702. Qwen3.8 Max (low) is sneller met 21.83s vs 32.78s, met slagingspercentages van 74.2% vs 80.3%.
Vergeleken modellen
- Rang
- #49
- Totaal aantal uitvoer-tokens
- 83,821
- Responstijd (gem.)
- 21.83s
- Totale kosten
- $0.702
Aanbevolen model
DeepSeek V4.1 Flash (high)
De score blijft dicht bij de beste score hier (8.5 vs 8.6) en het kost ongeveer 2.4x minder dan Qwen3.8 Max (low).
Gedetailleerde vergelijking
| Metriek | DeepSeek V4.1 Flash DeepSeek V4.1 Flash high | Qwen3.8 Max Qwen3.8 Max low |
|---|---|---|
| Score | 8.5 | 8.6 |
| Rang | #52 | #49 |
| Betrouwbaarheid | 9.6 | 10.0 |
| Consistentie | 8.9 | 9.0 |
| Pogingen | 66/66 | 66/66 |
| Correcte tests | ||
| Slaagpercentage per poging | 74.2% | 80.3% |
| Instabiele tests | 3 | 3 |
| Totaal runs | 66 | 66 |
| Kosten per resultaat | 1.912 | 4.384 |
| Totale kosten | $0.287 | $0.702 |
| Invoerprijs | $0.150 / 1M | $2.000 / 1M |
| Uitvoerprijs | $0.600 / 1M | $6.000 / 1M |
| Totaal aantal invoer-tokens | 100,111 | 99,181 |
| Uitvoer-tokens | 6,642 | 6,132 |
| Redeneer-tokens | 446,298 | 77,689 |
| Responstijd (gem.) | 32.78s | 21.83s |
| Responstijd (max) | 205.53s | 155.75s |
| Responstijd (totaal) | 721.15s | 480.37s |
| Parameters | 748B totaal (16B actief) | 2.4T totaal (~100B actief) |
| Beschikbaarheid | Open-source | Gesloten broncode |
Model generatie-showcase
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#52 DeepSeek V4.1 Flash
high- Kosten
- $0.035
- Tijd
- 111.0s
- Tokens
- 29,201 tok
#49 Qwen3.8 Max
low- Kosten
- $0.026
- Tijd
- 68.5s
- Tokens
- 4,280 tok
Topmodellen op score
Score vs totale kosten
Responstijd (gem.)
Score vs Responstijd (gem.)
Totaal aantal uitvoer-tokens
Score vs Totaal aantal uitvoer-tokens
Categorie-uitsplitsing
| Anti-AI-trucs | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 8.3 | 10.0 | 75.0% | 0 | 3.42s | 852 | 170 | 5,348 | |
| Qwen3.8 Max | 10.0 | 10.0 | 100.0% | 0 | 3.93s | 984 | 228 | 1,750 |
| Programmeren | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 49.19s | 7,509 | 376 | 108,285 | |
| Qwen3.8 Max | 8.4 | 7.4 | 88.9% | 1 | 28.69s | 8,127 | 512 | 15,952 |
| Gecombineerd | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 27.63s | 77,764 | 5,152 | 31,635 | |
| Qwen3.8 Max | 10.0 | 10.0 | 100.0% | 0 | 98.19s | 69,994 | 4,221 | 32,937 |
| Gegevensparsering en extractie | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 6.5 | 10.0 | 50.0% | 0 | 4.19s | 2,397 | 120 | 1,584 | |
| Qwen3.8 Max | 10.0 | 10.0 | 100.0% | 0 | 7.87s | 7,938 | 270 | 2,257 |
| Domeinspecifiek | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 5.3 | 7.2 | 44.4% | 1 | 127.30s | 909 | 28 | 249,623 | |
| Qwen3.8 Max | 2.9 | 7.2 | 11.1% | 1 | 40.12s | 1,014 | 62 | 19,512 |
| Algemene intelligentie | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 4.16s | 549 | 120 | 1,130 | |
| Qwen3.8 Max | 6.5 | 3.4 | 66.7% | 1 | 5.29s | 594 | 128 | 568 |
| Instructies opvolgen | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 7.06s | 783 | 63 | 1,972 | |
| Qwen3.8 Max | 10.0 | 10.0 | 100.0% | 0 | 4.01s | 855 | 102 | 851 |
| Puzzeloplossing | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 6.4 | 4.9 | 66.7% | 2 | 14.27s | 828 | 339 | 23,969 | |
| Qwen3.8 Max | 9.9 | 10.0 | 100.0% | 0 | 4.11s | 930 | 286 | 1,304 |
| Toolaanroepen | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 10.0 | 10.0 | 100.0% | 0 | 15.28s | 8,259 | 258 | 506 | |
| Qwen3.8 Max | 10.0 | 10.0 | 100.0% | 0 | 8.06s | 8,463 | 295 | 693 |
| Algemene kennis | Score | Consistentie | Slaagpercentage per poging | Instabiele tests | Correcte tests | Responstijd (gem.) | Invoer-tokens | Uitvoer-tokens | Redeneer-tokens |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 3.0 | 10.0 | 0.0% | 0 | 38.05s | 261 | 16 | 22,246 | |
| Qwen3.8 Max | 3.0 | 10.0 | 0.0% | 0 | 12.43s | 282 | 28 | 1,865 |
Snelle vergelijking
Vergelijkingspaar wisselen
DeepSeek V4.1 FlashhighvsGrok 4.5mediumDeepSeek V4.1 FlashhighvsMuse Spark 1.2lowGPT-5.4mediumvsQwen3.8 MaxlowDeepSeek V4.1 FlashhighvsGrok 4.6mediumDeepSeek V4.1 FlashhighvsGPT-5.4mediumClaude Fable 5mediumvsQwen3.8 MaxlowDeepSeek V4.1 FlashhighvsGrok 4.5lowMuse Spark 1.1mediumvsQwen3.8 MaxlowMuse Spark 1.2mediumvsQwen3.8 MaxlowDeepSeek V4.1 FlashhighvsGPT-5.2mediumDeepSeek V4.1 FlashmediumvsQwen3.8 MaxlowSeed 2.1 TurbomediumvsDeepSeek V4.1 Flashhigh