- Score
- 9.4 (+3.82 points)
- Total Cost (Current Price)
- $0.183 · 48.3% cheaper
- Response Time (avg)
- 6.24s · 78.1% faster
Summary
Mistral Large 4 scores 5.5 on AI BENCHY and ranks #274. It has 9.9 reliability, a 30.4% pass rate, $0.353 total cost, and 28.46s average response time.
Recommended alternatives
We recommend choosing one of these models instead:
- Score
- 7.5 (+2.00 points)
- Total Cost (Current Price)
- $0.000 · 100.0% cheaper
- Response Time (avg)
- 29.38s · 3.3% slower
- Score
- 8.3 (+2.72 points)
- Total Cost (Current Price)
- $0.047 · 86.7% cheaper
- Response Time (avg)
- 18.07s · 36.5% faster
Model facts
Researched on 2026-10-06
- Parameters
- 1.05T total (49B active)
- Architecture
- MoE
- Availability
- Closed
- License
- -
Mistral reports 1.05T total and 49B active parameters, plus a 1.6B vision encoder. The vendor describes this public preview as open-weight, but its checkpoint page currently supplies no weight download or license, and no Large 4 checkpoint was found in the official Hugging Face registry. Classified as API-only until the exact weights and license can be verified. The OpenRouter route exposes 524,288 context tokens and 262,144 completion tokens, with tools and structured outputs but no configurable reasoning parameter.
5.5
Consistency
8.8
9.9
$0.353
Total Output Tokens
84,218
Total Input Tokens
259,048
Input Price
$0.680 / 1M
Output Price
$2.090 / 1M
Cache Read Price
$0.070 / 1M
Cache Write Price
N/A
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Flaky tests
4
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Hamster playing table tennis
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
#274 Mistral Large 4
none- Cost
- $0.004
- Time
- 26.7s
- Tokens
- 1,923 tok
Charts
Choose the first model, then click a second model to open a side-by-side page.
Top Models by Score
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Quick Compare
Category Breakdown
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Agentic | 4.7 | 3.1 | |
| Anti-AI Tricks | 3.7 | 8.4 | |
| Coding | 5.4 | 7.8 | |
| Combined | 3.0 | 10.0 | |
| Data parsing and extraction | 10.0 | 10.0 | |
| Domain specific | 3.0 | 10.0 | |
| General Intelligence | 5.1 | 10.0 | |
| Instructions following | 6.5 | 10.0 | |
| Puzzle Solving | 6.1 | 7.1 | |
| Tool Calling | 10.0 | 10.0 | |
| Trivia | 3.0 | 10.0 |