AI BENCHY
AD
Track all your projects in one dashboard. Get 📊stats, 🔥heatmaps and 👀recordings in one self-hosted dashboard.
uxwizz.com

बदल नोंद

दिनांकानुसार गटबद्ध केलेली उत्पादन आणि बेंचमार्क अद्यतनांची साधी नोंद. आम्ही येथे नव्याने चाचणी केलेली मॉडेल्स, पुन्हा चाचण्या, बेंचमार्क बदल आणि प्रसिद्ध केलेले UX/उत्पादन काम नोंदवतो.

2026-09-05

  • नवीन वैशिष्ट्य: SVG उदाहरणे प्रतिमा म्हणून कॉपी करा किंवा वॉटरमार्क असलेले PNG डाउनलोड करा. पूर्वावलोकनात आता मॉडेलचा क्रमांक दिसतो.
  • बग दुरुस्ती: वैध SVG उदाहरणे पुन्हा उपलब्ध केली आणि वेळमर्यादा संपल्याचे संदेश स्पष्ट केले.

2026-09-04

  • नवीन चाचणी केलेली मॉडेल्स: GPT-6 Astra

2026-08-29

  • नवीन चाचणी केलेली मॉडेल्स: Hy4 preview

2026-08-26

  • नवीन चाचणी केलेली मॉडेल्स: GLM 5.3 Flash

2026-08-20

  • नवीन चाचणी केलेली मॉडेल्स: GLM 5.3

2026-08-15

  • नवीन वैशिष्ट्य: dots-studio/dots-3-note-preview:free उपलब्धता आणि रीझनिंग मोडच्या जलद चाचण्यांसाठी Dots3-Note Preview (मोफत) जोडले.
  • नवीन वैशिष्ट्य: स्थानिक रनच्या खर्च आणि मूल्य तुलनेत आता RTX 3090 चा अंदाजित वीज खर्च वापरला जातो.

2026-08-14

  • नवीन चाचणी केलेली मॉडेल्स: Gemini 3.7 Flash, Qwen3.8 27B
  • नवीन वैशिष्ट्य: Ollama द्वारे स्थानिक बेंचमार्क, प्रदात्याचा उगम, किंमत नसलेल्या स्थानिक संगणनाची योग्य हाताळणी आणि स्पष्ट स्थानिक रन टॅग जोडले.

2026-08-12

  • नवीन चाचणी केलेली मॉडेल्स: Seed 2.1 Turbo, Qwen3.8 2.4T A95B, Seed-2.0-Code, Grok 4.6
  • नवीन वैशिष्ट्य: स्रोतांसह पॅरामीटर संख्या, सक्रिय MoE पॅरामीटर्स, मॉडेलची उपलब्धता आणि आर्किटेक्चर यांची माहिती मॉडेल व तुलना पृष्ठांवर, तसेच लीडरबोर्डसाठी नवीन फिल्टर्स जोडली.

2026-08-08

  • नवीन चाचणी केलेली मॉडेल्स: Ling 3.0 Tiny

2026-08-06

  • नवीन चाचणी केलेली मॉडेल्स: Muse Spark 1.2

2026-08-04

  • नवीन चाचणी केलेली मॉडेल्स: Qwen3.8 Max

2026-08-02

  • नवीन वैशिष्ट्य: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.

2026-07-28

  • नवीन चाचणी केलेली मॉडेल्स: Trinity Large Thinking, Qwen3.7 Flash
  • नवीन वैशिष्ट्य: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.

2026-07-26

  • नवीन वैशिष्ट्य: Added editorial model recommendations for affordable, intelligent, and fast choices.
  • UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.

2026-07-25

  • नवीन चाचणी केलेली मॉडेल्स: Claude Opus 5

2026-07-24

  • नवीन चाचणी केलेली मॉडेल्स: Ling-3.0-flash

2026-07-22

  • नवीन चाचणी केलेली मॉडेल्स: Laguna S 2.1

2026-07-20

  • नवीन चाचणी केलेली मॉडेल्स: LongCat 2.0

2026-07-18

  • नवीन चाचणी केलेली मॉडेल्स: Inkling

2026-07-17

  • नवीन चाचण्या जोडल्या: Added an executable JavaScript tool-calling benchmark with reference-case validation.

2026-07-16

  • नवीन चाचणी केलेली मॉडेल्स: Muse Spark 1.1, Kimi K3
  • नवीन वैशिष्ट्य: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.

2026-07-03

  • नवीन वैशिष्ट्य: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.

2026-07-02

  • नवीन चाचणी केलेली मॉडेल्स: Laguna XS 2.1

2026-06-30

  • नवीन चाचणी केलेली मॉडेल्स: Claude Sonnet 5

2026-06-18

  • नवीन चाचणी केलेली मॉडेल्स: North Mini Code

2026-06-17

  • नवीन चाचणी केलेली मॉडेल्स: GLM 5.2
  • बग दुरुस्ती: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
  • UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.

2026-06-16

  • नवीन वैशिष्ट्य: Added cost sorting and filtering across leaderboard and category views.

2026-06-12

  • नवीन चाचणी केलेली मॉडेल्स: Kimi K2.7 Code
  • नवीन वैशिष्ट्य: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.

2026-06-10

  • नवीन चाचणी केलेली मॉडेल्स: Claude Fable 5

2026-06-06

  • नवीन वैशिष्ट्य: Model showcases now support shareable lightbox views, filtering, and score details.

2026-06-05

  • नवीन वैशिष्ट्य: Added model-generated visual showcases to model and comparison pages.
  • नवीन वैशिष्ट्य: Added multi-model comparison pages with model recommendations and category-based ranking.

2026-06-04

  • नवीन चाचणी केलेली मॉडेल्स: Nemotron 3 Ultra
  • नवीन चाचण्या जोडल्या: Added a coding benchmark with executable reference cases.

2026-06-03

  • नवीन चाचणी केलेली मॉडेल्स: Qwen3.7 Plus

2026-06-01

  • नवीन चाचणी केलेली मॉडेल्स: MiniMax M3

2026-05-29

  • नवीन चाचणी केलेली मॉडेल्स: Step 3.7 Flash

2026-05-28

  • नवीन चाचणी केलेली मॉडेल्स: Claude Opus 4.8

2026-05-27

  • नवीन वैशिष्ट्य: Added current-price cost calculations while preserving original tested-at pricing for auditability.

2026-05-22

  • नवीन चाचणी केलेली मॉडेल्स: Qwen3.7 Max
  • नवीन चाचण्या जोडल्या: C++ सोल्यूशनमध्ये बग शोधण्यावर केंद्रित असलेली नवीन Coding चाचणी श्रेणी जोडली.

2026-05-21

  • नवीन चाचणी केलेली मॉडेल्स: Grok Build 0.1
  • नवीन चाचण्या जोडल्या: Added a new benchmark test and improved answer judging.
  • बग दुरुस्ती: प्रोव्हायडर पडताळणीत reasoning आवश्यक ठरल्यानंतर xAI Grok Build 0.1 चा असमर्थित no-reasoning प्रकार काढला.

2026-05-20

  • नवीन चाचणी केलेली मॉडेल्स: Gemini 3.5 Flash

2026-05-10

  • नवीन चाचणी केलेली मॉडेल्स: Ring-2.6-1T

2026-05-08

  • नवीन चाचण्या जोडल्या: Added a new benchmark test to expand suite coverage.
  • बग दुरुस्ती: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto.
  • UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.

2026-05-06

  • नवीन चाचणी केलेली मॉडेल्स: Cobuddy

2026-04-30

  • नवीन चाचणी केलेली मॉडेल्स: Owl Alpha

2026-04-26

  • UX: मोबाइलवरील तुलना ड्रॉपडाउनची जागा सुधारली, मॉडेल पेज लेआउट अधिक घट्ट केले, आणि रन इतिहास प्रति-मॉडेल शार्डमध्ये विभागला जेणेकरून पेज कमी ऐतिहासिक डेटा लोड करतील.
  • बग दुरुस्ती: रन इतिहास आता त्याच suite मधील जवळपास-डुप्लिकेट री-टेस्ट गटबद्ध करतो आणि मॉडेल पृष्ठांवर सर्व सार्वजनिक रन थेट तुलना तक्त्यात दाखवतो.

2026-04-25

  • नवीन वैशिष्ट्य: विश्वसनीयता स्कोअर टेलीमेट्री जोडली, त्यामुळे लक्ष्य API आणि रेट-लिमिट अपयशे चुकीच्या उत्तरांपासून वेगळी ट्रॅक होतात.

2026-04-24

  • नवीन चाचणी केलेली मॉडेल्स: DeepSeek V4 Flash 0423, DeepSeek V4 Pro, GPT-5.5
  • बग दुरुस्ती: Changelog मधील model links आता canonical live model pages कडे जातात, आणि model pages आता reasoning variants मध्ये परस्पर links देतात.

2026-04-23

  • नवीन चाचणी केलेली मॉडेल्स: Ling-2.6-1T, Hy3 preview
  • नवीन वैशिष्ट्य: रन इतिहास - मॉडेल पृष्ठे आता ऐतिहासिक public runs आणि side-by-side run comparison तक्ता दाखवतात. (उदाहरण मॉडेल पृष्ठ)
  • UX: लीडरबोर्ड आता URL-आधारित pagination, filters आणि ranking list मधून direct compare actions ला support करतो.
  • बग दुरुस्ती: होमपेज search, filter counts आणि pagination state आता संपूर्ण dataset मध्ये सुसंगत राहतात.
  • पुन्हा चाचणी: GLM 5.1 या मॉडेलसाठी पूर्ण benchmark suite पुन्हा चालवली आणि public run-history snapshot स्वच्छ केला.
  • बग दुरुस्ती: ज्या मॉडेल्सचा प्रत्यक्ष retest झाला नाही त्यांना नवीन tested_at timestamp मिळू नये असे केले.

2026-04-20

  • नवीन चाचणी केलेली मॉडेल्स: Kimi K2.6

2026-04-16

  • नवीन चाचणी केलेली मॉडेल्स: Claude Opus 4.7

2026-04-14

  • नवीन चाचणी केलेली मॉडेल्स: Ling-2.6-flash

2026-04-11

  • नवीन चाचणी केलेली मॉडेल्स: GLM 5.1

2026-04-04

  • नवीन चाचणी केलेली मॉडेल्स: Gemma 4 26B A4B

2026-03-21

  • नवीन चाचणी केलेली मॉडेल्स: Mimo V2 Omni

2026-03-20

  • नवीन चाचणी केलेली मॉडेल्स: Mimo V2 PRO

2026-03-18

  • नवीन चाचणी केलेली मॉडेल्स: MiniMax M2.7
  • नवीन वैशिष्ट्य: Added input and output pricing metrics to model and comparison pages.

2026-03-15

  • नवीन चाचणी केलेली मॉडेल्स: GLM 5 Turbo

2026-03-12

2026-03-06

  • नवीन वैशिष्ट्य: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.

2026-03-05

  • नवीन चाचणी केलेली मॉडेल्स: Mercury 2, GPT-5.4

2026-03-04

  • नवीन चाचणी केलेली मॉडेल्स: GPT-5.2 Chat, GPT 5.3 Chat
  • नवीन वैशिष्ट्य: Added interactive comparison charts for direct model analysis.
  • नवीन चाचण्या जोडल्या: Added a VAT compliance micro-audit benchmark with structured tool requirements.

2026-03-02

  • नवीन वैशिष्ट्य: Added share actions with copyable links across model and comparison pages.

2026-02-27

  • नवीन चाचणी केलेली मॉडेल्स: Seed-2.0-Mini
  • नवीन वैशिष्ट्य: Added reasoning-quality and consistency charts to model comparisons.

2026-02-24

  • नवीन चाचणी केलेली मॉडेल्स: GPT-5.3-Codex

2026-02-17

  • नवीन चाचणी केलेली मॉडेल्स: Claude Sonnet 4.6

चेंजलॉग पृष्ठ तयार केले

हा चेंजलॉग लाँचनंतर सुरू झाला, त्यामुळे काही जुनी अद्यतने येथे नाहीत.