The average score is effectively tied at 1.0 vs 1.0. Tev1 4B Experimental has the lower benchmark cost at $0.002 vs $0.009. Tev1 4B Experimental is faster at 498ms vs 6.01s, with pass rates of 13.0% vs 14.5%.
Last updated at:
2026-10-01
Compared models
Rank
#381
Total Output Tokens
96
Response Time (avg)
498ms
Total Cost
$0.002
Rank
#380
Total Output Tokens
51
Response Time (avg)
6.01s
Total Cost
$0.009
Recommended modelTev1 4B Experimental
It has the best score here (1.0), while costing about 6.0x less than Solar Decide.
Detailed comparison
Metric
Tev1 4B ExperimentalTev1 4B Experimentalnone21/69 attempts. Choices supplied by adapters. Supports 7/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-30
Solar DecideSolar Decidenone21/69 attempts. Choices supplied by adapters. Supports 7/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-28
Metric
Tev1 4B ExperimentalTev1 4B Experimentalnone21/69 attempts. Choices supplied by adapters. Supports 7/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-30
Solar DecideSolar Decidenone21/69 attempts. Choices supplied by adapters. Supports 7/23 tests. Other tests are unsupported, not failed. Score uses the full suite.Release: 2026-09-28
Score
1.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
Rank
#381
#380
Reliability
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
10.0First-attempt success score: 10.0 means no retryable target API or rate-limit failures before successful calls; tracked failures lower the score.…
Consistency
3.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
2.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
Attempts
21/69
21/69
Tests Correct
A test is fully passed only if every run passed for that test.Wrong answer: 4Response Time (avg)498msResponse Time (max)602msResponse Time (total)3.49sA test is fully passed only if every run passed for that test.…
A test is fully passed only if every run passed for that test.Wrong answer: 5Response Time (avg)6.01sResponse Time (max)13.10sResponse Time (total)42.05sA test is fully passed only if every run passed for that test.…
Attempt pass rate
13.0%Attempt pass rate = passed attempts / total attempts across runs.…
14.5%Attempt pass rate = passed attempts / total attempts across runs.…
Flaky tests
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
3Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
Total Runs
21Total Runs…
21Total Runs…
Cost per result
0.048
0.119
Total Cost
$0.002
$0.009
Input Price
$0.042 / 1MInput Price…
$0.050 / 1MInput Price…
Output Price
$0.000 / 1MOutput Price…
$0.000 / 1MOutput Price…
Cache Read Price
N/ACache Read Price…
$0.050 / 1MCache Read Price…
Cache Write Price
N/ACache Write Price…
N/ACache Write Price…
Total Input Tokens
33,864Total Input Tokens…
47,586Total Input Tokens…
Output Tokens
96Output Tokens…
51Output Tokens…
Reasoning Tokens
0Reasoning Tokens…
0Reasoning Tokens…
Response Time (avg)
498msResponse Time (avg)…
6.01sResponse Time (avg)…
Response Time (max)
602msResponse Time (max)…
13.10sResponse Time (max)…
Response Time (total)
3.49sResponse Time (total)…
42.05sResponse Time (total)…
Parameters
4B
~35B total (~3B active)
Availability
Weights available
Closed
Cache prices apply to input tokens. Reads reuse cached prompts; writes store them and can cost extra. Output tokens use the output price.
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
2.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
2.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
25.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)602msResponse Time (max)602msResponse Time (total)602msA test is fully passed only if every run passed for that test.…
0.8Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
2.5Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)2.41sResponse Time (max)2.41sResponse Time (total)2.41sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
3.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
10.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)509msResponse Time (max)563msResponse Time (total)1.02sA test is fully passed only if every run passed for that test.…
3.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
1.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
2Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 2Response Time (avg)5.00sResponse Time (max)6.86sResponse Time (total)10.00sA test is fully passed only if every run passed for that test.…
6.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
6.7Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
66.7%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)440msResponse Time (max)491msResponse Time (total)879msA test is fully passed only if every run passed for that test.…
4.3Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.9Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
44.4%Attempt pass rate = passed attempts / total attempts across runs.…
1Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)7.26sResponse Time (max)13.10sResponse Time (total)14.52sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
1.5Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)555msResponse Time (max)555msResponse Time (total)555msA test is fully passed only if every run passed for that test.…
5.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
5.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
50.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)11.66sResponse Time (max)11.66sResponse Time (total)11.66sA test is fully passed only if every run passed for that test.…
1.7Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)432msResponse Time (max)432msResponse Time (total)432msA test is fully passed only if every run passed for that test.…
1.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
3.3Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.Wrong answer: 1Response Time (avg)3.47sResponse Time (max)3.47sResponse Time (total)3.47sA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…
0.0Summarizes broad quality across our full private benchmark suite, so ranking reflects consistent performance.…
0.0Consistency score reflects run-to-run stability (10 = very consistent, even if consistently wrong).…
0.0%Attempt pass rate = passed attempts / total attempts across runs.…
0Flaky tests had mixed outcomes across runs (at least one pass and one fail).…
A test is fully passed only if every run passed for that test.No failed answers.Response Time (avg)0msResponse Time (max)0msResponse Time (total)0msA test is fully passed only if every run passed for that test.…