#217 Hy4 preview
low- Cost
- $0.000
- Time
- 1.1s
- Tokens
- 0 tok
Summary
Hy4 preview scores 5.6 on AI BENCHY and ranks #217. It has 5.4 reliability, a 19.7% pass rate, $1.745 total cost, and 214.25s average response time.
What makes Hy4 preview unique: It stands out most in General Intelligence, where it ranks #1, while Combined is its weakest area at #11. It uses unusually many reasoning tokens, which can help explain its slower or more expensive runs.
Researched on 2026-08-29
Vendor-reported backbone counts exclude the native 10B total / 0.7B active MTP layer used for speculative decoding.
5.6
Consistency
8.5
5.4
$1.745
Total Output Tokens
682,538
Total Input Tokens
45,154
Input Price
$0.834 / 1M
Output Price
$2.501 / 1M
Flaky tests
3
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Prompt: Create a detailed SVG illustration of a hamster playing table tennis.
Run history
| Tested on | Score | Reliability | Tests Correct | Total Cost | Compare |
|---|---|---|---|---|---|
| 2026-08-29 15:57 Re-test | 5.1 | 3.6 | $1.681 | Current run | |
| 2026-08-29 14:28 Initial run | 4.3 | 1.6 | $1.433 | Compare |
Choose the first model, then click a second model to open a side-by-side page.
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Anti-AI Tricks | 6.0 | 8.9 | |
| Coding | 5.2 | 8.6 | |
| Combined | 2.9 | 5.8 | |
| Data parsing and extraction | 8.7 | 10.0 | |
| Domain specific | 4.5 | 10.0 | |
| General Intelligence | 7.5 | 10.0 | |
| Instructions following | 9.8 | 10.0 | |
| Puzzle Solving | 5.3 | 5.0 | |
| Tool Calling | 3.0 | 10.0 | |
| Trivia | 3.0 | 10.0 |