Summary
Tev1 4B Experimental scores 1.0 on AI BENCHY and ranks #381. It has 10.0 reliability, a 13.0% pass rate, $0.002 total cost, and 498ms average response time.
What makes Tev1 4B Experimental unique: Its total benchmark cost is unusually low for its score range. It is notably fast compared with similar models.
Model facts
Researched on 2026-10-01
- Parameters
- 4B
- Architecture
- Dense
- Availability
- Weights available
- License
- Pending; see model card
Best estimate from public evidence; the vendor did not disclose every value. Vendor-designated 4B Qwen3.5-4B supervised fine-tune with the standard next-token head. Runnable training/inference code is MIT; the exact fine-tuned weight release license is still being finalized. Uses one option letter with thinking disabled, not Jev architecture. The 4B designation is rounded, not an exact combined parameter count.
1.0
Consistency
3.0
10.0
$0.002
Total Output Tokens
96
Total Input Tokens
33,864
Input Price
$0.042 / 1M
Output Price
$0.000 / 1M
Cache Read Price
N/A
Cache Write Price
N/A
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
Wrong Tests: 4
Attempt pass rate: 13.0%
Benchmark coverage: 21/69 attempts. Choices supplied by adapters. Supports 7/23 tests. Other tests are unsupported, not failed. Score uses the full suite.
Flaky tests
0
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Charts
Choose the first model, then click a second model to open a side-by-side page.
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Quick Compare
Category Breakdown
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Agentic | 0.0 | 0.0 | |
| Anti-AI Tricks | 2.5 | 2.5 | |
| Coding | 0.0 | 0.0 | |
| Combined | 0.0 | 0.0 | |
| Data parsing and extraction | 3.0 | 10.0 | |
| Domain specific | 6.7 | 6.7 | |
| General Intelligence | 0.0 | 0.0 | |
| Instructions following | 1.5 | 5.0 | |
| Puzzle Solving | 1.7 | 3.3 | |
| Tool Calling | 0.0 | 0.0 | |
| Trivia | 0.0 | 0.0 |