Urambazaji
AI BENCHY
Linganisha Chati
❤️ Made by XCS
Your ad here

AI BENCHY Compare

OpenAI: GPT-5.3-Codex vs StepFun: Step 3.5 Flash

Jina la modeli:

Benchmark zimetengenezwa kutoka seti za majaribio za AI BENCHY tarehe : 2026-02-27 15:16

Muhtasari

Kipimo OpenAI: GPT-5.3-Codex medium Toleo: Tarehe ya kutolewa haijulikani StepFun: Step 3.5 Flash medium Toleo: Tarehe ya kutolewa haijulikani Inapatikana bure
Nafasi #7 #11
Alama 7.93 7.00
Uthabiti 8.84 8.32
Gharama kwa matokeo 4.641 0.000
Jumla ya gharama $0.465 $0.000
Majaribio sahihi
Majaribio yenye makosa 4 5
Kiwango cha kupita kwa kila jaribio 78.6% 73.8%
Majaribio yasiyo thabiti 2 3
Tokeni za matokeo 1,201 60,502
Tokeni za hoja 30,056 117,044

Mgawanyo wa kategoria

Mbinu za kupinga AI Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 10.00 10.00 100.0% 0 216 1,421
StepFun: Step 3.5 Flash 10.00 10.00 100.0% 0 13,924 17,208
Uchanganuzi na uchimbaji wa data Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 10.00 10.00 100.0% 0 234 735
StepFun: Step 3.5 Flash 10.00 10.00 100.0% 0 535 11,548
Mahususi kwa domeni Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 4.00 7.21 55.6% 1 64 25,308
StepFun: Step 3.5 Flash 4.00 7.21 44.4% 1 40,942 74,237
Ufuataji wa maagizo Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 9.00 10.00 50.0% 0 93 693
StepFun: Step 3.5 Flash 10.00 10.00 100.0% 0 2,121 3,274
Puzzle Solving Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 7.00 7.38 77.8% 1 340 1,407
StepFun: Step 3.5 Flash 2.00 4.96 33.3% 2 2,705 6,975
Mwito wa zana Alama Uthabiti Kiwango cha kupita kwa kila jaribio Majaribio yasiyo thabiti Majaribio sahihi Tokeni za matokeo Tokeni za hoja
OpenAI: GPT-5.3-Codex 10.00 10.00 100.0% 0 254 492
StepFun: Step 3.5 Flash 10.00 10.00 100.0% 0 275 3,802

Badilisha jozi ya ulinganisho