#119 MiniMax M2.5
medium- Cost
- $0.000
- Time
- 300.0s
- Tokens
- 0 tok
Summary
MiniMax M2.5 scores 5.4 on AI BENCHY and ranks #119. It has 8.3 reliability, a 50.0% pass rate, $0.305 total cost, and 50.25s average response time.
Researched on 2026-08-12
5.4
Consistency
6.1
8.3
$0.305
Total Output Tokens
360,672
Total Input Tokens
0
Input Price
$0.150 / 1M
Output Price
$1.150 / 1M
Cache Read Price
N/A
Cache Write Price
N/A
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Flaky tests
10
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
Run history
| Tested on | Score | Reliability | Tests Correct | Total Cost | Compare |
|---|---|---|---|---|---|
| 2026-08-14 01:38 Re-test | 4.6 | 10.0 | $0.272 ↓ | Compare | |
| 2026-07-16 21:18 New test added | 4.6 | 10.0 | $0.340 ↓ | Compare | |
| 2026-06-04 13:23 New test added | 5.3 | 8.9 | $0.385 ↓ | Compare | |
| 2026-05-21 23:48 Suite changed | 5.4 | 8.3 | $0.305 | Current run | |
| 2026-04-20 17:48 First recorded run | 5.7 | N/A | $0.250 | Compare |
This run used a different benchmark suite. Keep suite changes in mind when reading historical movement.
Choose the first model, then click a second model to open a side-by-side page.
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Anti-AI Tricks | 7.9 | 6.3 | |
| Coding | 3.5 | 9.8 | |
| Combined | 4.5 | 2.1 | |
| Data parsing and extraction | 4.6 | 1.7 | |
| Domain specific | 2.9 | 4.4 | |
| General Intelligence | 3.8 | 2.5 | |
| Instructions following | 7.5 | 6.7 | |
| Puzzle Solving | 5.3 | 7.2 | |
| Tool Calling | 10.0 | 10.0 | |
| Trivia | 3.0 | 10.0 |