AI BENCHY
Advertise here

تبدیلی لاگ

مصنوعہ اور بینچ مارک اپڈیٹس کا ایک سادہ لاگ، تاریخ کے مطابق گروپ کیا گیا۔ ہم اسے نئے ٹیسٹ کیے گئے ماڈلز، دوبارہ ٹیسٹ، بینچ مارک تبدیلیوں، اور جاری کی گئی UX/پروڈکٹ اپڈیٹس کو نوٹ کرنے کے لیے استعمال کرتے ہیں۔

2026-09-05

  • نئی خصوصیت: SVG نمونے تصاویر کے طور پر کاپی کریں یا واٹرمارک والے PNG ڈاؤن لوڈ کریں۔ پیش منظر میں اب ماڈل کا درجہ دکھایا جاتا ہے۔
  • بگ فکس: درست SVG نمونے بحال کیے گئے اور وقت کی حد ختم ہونے کی خرابیاں واضح کی گئیں۔

2026-09-04

  • نئے ٹیسٹ کیے گئے ماڈلز: GPT-6 Astra

2026-08-29

  • نئے ٹیسٹ کیے گئے ماڈلز: Hy4 preview

2026-08-26

2026-08-20

  • نئے ٹیسٹ کیے گئے ماڈلز: GLM 5.3

2026-08-15

  • نئی خصوصیت: dots-studio/dots-3-note-preview:free دستیابی اور ریزننگ موڈز کی مختصر جانچ کے لیے Dots3-Note Preview (مفت) شامل کیا گیا۔
  • نئی خصوصیت: مقامی رنز کی لاگت اور قدر کے موازنے میں اب RTX 3090 کی تخمینی بجلی لاگت استعمال ہوتی ہے۔

2026-08-14

  • نئے ٹیسٹ کیے گئے ماڈلز: Gemini 3.7 Flash, Qwen3.8 27B
  • نئی خصوصیت: Ollama کے ذریعے مقامی بینچ مارک اجرا، فراہم کنندہ کی معلومات، بغیر قیمت والی مقامی کمپیوٹنگ کی درست ہینڈلنگ اور واضح مقامی رن ٹیگ شامل کیے گئے۔

2026-08-12

  • نئے ٹیسٹ کیے گئے ماڈلز: Seed 2.1 Turbo, Qwen3.8 2.4T A95B, Seed-2.0-Code, Grok 4.6
  • نئی خصوصیت: ماڈل اور موازنہ صفحات میں حوالوں کے ساتھ پیرامیٹرز کی تعداد، فعال MoE پیرامیٹرز، ماڈل کی دستیابی اور ساخت شامل کی گئی، اور لیڈر بورڈ میں نئے فلٹرز شامل کیے گئے۔

2026-08-08

2026-08-04

  • نئے ٹیسٹ کیے گئے ماڈلز: Qwen3.8 Max

2026-08-02

  • نئی خصوصیت: Model pages now show benchmark-suite coverage so incomplete results are clear at a glance.

2026-07-28

  • نئے ٹیسٹ کیے گئے ماڈلز: Trinity Large Thinking, Qwen3.7 Flash
  • نئی خصوصیت: Added site-wide Spotlight Search with keyboard navigation and a shortcut for finding pages, models, and comparisons.

2026-07-26

  • نئی خصوصیت: Added editorial model recommendations for affordable, intelligent, and fast choices.
  • UX: Comparison pages now use draggable model summary cards with clearer model selection and ordering.

2026-07-25

2026-07-22

2026-07-20

  • نئے ٹیسٹ کیے گئے ماڈلز: LongCat 2.0

2026-07-18

  • نئے ٹیسٹ کیے گئے ماڈلز: Inkling

2026-07-17

  • نئے ٹیسٹ شامل کیے گئے: Added an executable JavaScript tool-calling benchmark with reference-case validation.

2026-07-16

  • نئے ٹیسٹ کیے گئے ماڈلز: Muse Spark 1.1, Kimi K3
  • نئی خصوصیت: The benchmark runner now supports executable tools with validated tool-call schemas and sandboxed execution.

2026-07-03

  • نئی خصوصیت: Launched AI World Cup, where benchmarked models generate football strategies and compete in simulated matches.

2026-07-02

2026-06-17

  • نئے ٹیسٹ کیے گئے ماڈلز: GLM 5.2
  • بگ فکس: Adjusted missing-test handling so models are not scored as if unavailable tests were valid wrong answers.
  • UX: Leaderboard search now supports comma-separated model queries, so searches like "deepseek, glm" show matches for either model family.

2026-06-16

  • نئی خصوصیت: Added cost sorting and filtering across leaderboard and category views.

2026-06-12

  • نئے ٹیسٹ کیے گئے ماڈلز: Kimi K2.7 Code
  • نئی خصوصیت: Updated scoring to use per-category bias adjustments, so category-level differences are normalized before they roll into leaderboard results.

2026-06-06

  • نئی خصوصیت: Model showcases now support shareable lightbox views, filtering, and score details.

2026-06-05

  • نئی خصوصیت: Added model-generated visual showcases to model and comparison pages.
  • نئی خصوصیت: Added multi-model comparison pages with model recommendations and category-based ranking.

2026-06-04

  • نئے ٹیسٹ کیے گئے ماڈلز: Nemotron 3 Ultra
  • نئے ٹیسٹ شامل کیے گئے: Added a coding benchmark with executable reference cases.

2026-06-03

2026-06-01

  • نئے ٹیسٹ کیے گئے ماڈلز: MiniMax M3

2026-05-27

  • نئی خصوصیت: Added current-price cost calculations while preserving original tested-at pricing for auditability.

2026-05-22

  • نئے ٹیسٹ کیے گئے ماڈلز: Qwen3.7 Max
  • نئے ٹیسٹ شامل کیے گئے: C++ کے حل میں کیڑے تلاش کرنے پر مرکوز ایک نیا Coding ٹیسٹ زمرہ شامل کیا گیا۔

2026-05-21

  • نئے ٹیسٹ کیے گئے ماڈلز: Grok Build 0.1
  • نئے ٹیسٹ شامل کیے گئے: Added a new benchmark test and improved answer judging.
  • بگ فکس: Provider validation کی طرف سے reasoning لازم قرار دینے کے بعد xAI Grok Build 0.1 کا غیر معاون no-reasoning ویریئنٹ ہٹا دیا گیا۔

2026-05-10

  • نئے ٹیسٹ کیے گئے ماڈلز: Ring-2.6-1T

2026-05-08

  • نئے ٹیسٹ شامل کیے گئے: Added a new benchmark test to expand suite coverage.
  • بگ فکس: Reasoning chips and compare labels now recognize the minimal reasoning variant instead of falling back to auto.
  • UX: Model pages now order sibling reasoning-variant chips from highest effort to lowest.

2026-05-06

  • نئے ٹیسٹ کیے گئے ماڈلز: Cobuddy

2026-04-30

  • نئے ٹیسٹ کیے گئے ماڈلز: Owl Alpha

2026-04-26

  • UX: موبائل پر موازنہ ڈراپ ڈاؤن کی جگہ بہتر کی، ماڈل پیج لے آؤٹ کو زیادہ مختصر کیا، اور رن ہسٹری کو ہر ماڈل کے شاردز میں تقسیم کیا تاکہ صفحات کم تاریخی ڈیٹا لوڈ کریں۔
  • بگ فکس: رن ہسٹری اب اسی suite کے قریباً ڈپلیکیٹ ری ٹیسٹس کو گروپ کرتی ہے اور ماڈل صفحات پر تمام عوامی رنز کو براہ راست موازنہ جدول میں دکھاتی ہے۔

2026-04-25

  • نئی خصوصیت: اعتماد پذیری اسکور ٹیلیمیٹری شامل کی گئی تاکہ ہدف API اور ریٹ لمٹ ناکامیاں غلط جوابات سے الگ ٹریک ہوں۔

2026-04-24

  • نئے ٹیسٹ کیے گئے ماڈلز: DeepSeek V4 Flash 0423, DeepSeek V4 Pro, GPT-5.5
  • بگ فکس: چینج لاگ میں ماڈل لنکس اب معیاری لائیو ماڈل صفحات پر جاتے ہیں، اور ماڈل صفحات اب reasoning variants کے درمیان بھی لنک دیتے ہیں۔

2026-04-23

  • نئے ٹیسٹ کیے گئے ماڈلز: Ling-2.6-1T, Hy3 preview
  • نئی خصوصیت: رن ہسٹری - ماڈل صفحات اب تاریخی public runs اور side-by-side run comparison جدول دکھاتے ہیں۔ (مثالی ماڈل صفحہ)
  • UX: لیڈر بورڈ اب URL-based pagination، filters اور ranking list سے direct compare actions کو support کرتا ہے۔
  • بگ فکس: ہوم پیج search، filter counts اور pagination state اب پورے dataset میں ایک جیسی رہتی ہے۔
  • دوبارہ ٹیسٹ: GLM 5.1 اس ماڈل کے لیے مکمل benchmark suite دوبارہ چلائی گئی اور public run-history snapshot صاف کیا گیا۔
  • بگ فکس: جن ماڈلز کا حقیقت میں retest نہیں ہوا، انہیں نیا tested_at timestamp ملنا بند کر دیا گیا۔

2026-04-20

  • نئے ٹیسٹ کیے گئے ماڈلز: Kimi K2.6

2026-04-11

  • نئے ٹیسٹ کیے گئے ماڈلز: GLM 5.1

2026-03-21

2026-03-20

  • نئے ٹیسٹ کیے گئے ماڈلز: Mimo V2 PRO

2026-03-18

  • نئے ٹیسٹ کیے گئے ماڈلز: MiniMax M2.7
  • نئی خصوصیت: Added input and output pricing metrics to model and comparison pages.

2026-03-15

  • نئے ٹیسٹ کیے گئے ماڈلز: GLM 5 Turbo

2026-03-06

  • نئی خصوصیت: Added the Methodology section and expanded charts with latency, cost, token, and model-switching views.

2026-03-04

  • نئے ٹیسٹ کیے گئے ماڈلز: GPT-5.2 Chat, GPT 5.3 Chat
  • نئی خصوصیت: Added interactive comparison charts for direct model analysis.
  • نئے ٹیسٹ شامل کیے گئے: Added a VAT compliance micro-audit benchmark with structured tool requirements.

2026-03-02

  • نئی خصوصیت: Added share actions with copyable links across model and comparison pages.

2026-02-27

  • نئے ٹیسٹ کیے گئے ماڈلز: Seed-2.0-Mini
  • نئی خصوصیت: Added reasoning-quality and consistency charts to model comparisons.

2026-02-24

چینج لاگ صفحہ بنایا گیا

یہ چینج لاگ لانچ کے بعد شروع ہوا، اس لیے کچھ پرانی اپڈیٹس یہاں موجود نہیں ہیں۔