Summary
Kev 4B scores 3.0 on AI BENCHY and ranks #366. It has 10.0 reliability, a 16.7% pass rate, $0.002 total cost, and 5.63s average response time.
What makes Kev 4B unique: Its total benchmark cost is unusually low for its score range. It is notably fast compared with similar models.
Model facts
Researched on 2026-09-28
- Parameters
- 4B
- Architecture
- Dense
- Availability
- Open source
- License
- Apache-2.0
Best estimate from public evidence; the vendor did not disclose every value. The current model card reports a 4B Qwen3.5-4B-Base backbone with a 33.8M-parameter LoRA adapter and pointer head. The rounded 4B size is the model designation, not an exact combined parameter count. Dense architecture is inferred from the base checkpoint. Apache-2.0 weights, adapter/head, and runnable training/serving code are public. This checkpoint returns typed decisions, not generated text.
3.0
Consistency
5.1
10.0
$0.002
Total Output Tokens
28,061
Total Input Tokens
42,804
Input Price
$0.042 / 1M
Output Price
$0.000 / 1M
Wrong Tests: 9
Attempt pass rate: 16.7%
Benchmark coverage: 36/66 attempts. Choices supplied by adapters. Supports 12/22 tests. Other tests are unsupported, not failed. Score uses the full suite.
Flaky tests
1
Flaky tests had mixed outcomes across runs (at least one pass and one fail).
Charts
Choose the first model, then click a second model to open a side-by-side page.
Score vs Total Cost
Response Time (avg)
Score vs Response Time (avg)
Total Output Tokens
Score vs Total Output Tokens
Quick Compare
Category Breakdown
| Category | Score | Consistency | Tests Correct |
|---|---|---|---|
| Anti-AI Tricks | 2.3 | 7.5 | |
| Coding | 4.5 | 6.7 | |
| Combined | 0.0 | 0.0 | |
| Data parsing and extraction | 6.5 | 10.0 | |
| Domain specific | 5.3 | 10.0 | |
| General Intelligence | 0.0 | 0.0 | |
| Instructions following | 1.5 | 5.0 | |
| Puzzle Solving | 2.0 | 10.0 | |
| Tool Calling | 0.0 | 0.0 | |
| Trivia | 0.0 | 0.0 |