Da7em Benchداحم بنش
An independent benchmark for AI models, built on real client work. The first of its kind in the world.
How it works:
Every model runs about 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication.
Each model is tested across several harnesses, both official and neutral ones (Droid, Hermes Agent, Devin, Cursor), so no single harness decides a model's fate and the results reflect the model.
Scoring is 1 to 5. A 5 means the work was accepted as delivered. Middle scores mean it needed revision. A 1 means it failed.
The bar is professional work. Every result is judged against what a paid professional would have delivered for the same brief.
Tasks stay private so they can't leak into training data and inflate future scores.
The goal isn't one more leaderboard. It's helping you pick the right model for your kind of work. A model that leads in reasoning can still fall behind in writing or design taste, and the radar charts show exactly where.
This is v0.1. It will keep evolving with harder, market-relevant tasks and with new models as they ship. A few popular models (Opus 5, GPT Luna) aren't included yet because I haven't run enough tasks on them to score them fairly.
Da7em Bench is fully independent. No sponsors, no vendor relationships. I've paid for every run out of pocket, thousands of dollars so far. That's the whole point: an honest, neutral look at what these models actually do on real work.
معيار مستقل لنماذج الذكاء الاصطناعي، بُني على أعمال حقيقية لعملاء. وهو الأول من نوعه في العالم.
كيف يعمل:
يخوض كل نموذج نحو ٢٠٠ مهمة حقيقية في كل واحد من ١٢ مجالًا: الاستدلال، والبحث، والتخطيط، والإنجاز، والمثابرة، والدقة، والصدق، والقبول، والهندسة، والذائقة، والكتابة، والتواصل.
يُختبر كل نموذج عبر عدة بيئات تشغيل، رسمية ومحايدة (درويد، هيرميس إيجنت، ديفن، كيرسر)، حتى لا تحدد بيئة واحدة مصيره وتعكس النتائج أداء النموذج نفسه.
الدرجات من ١ إلى ٥. تعني ٥ أن العمل قُبل كما سُلّم. وتعني الدرجات المتوسطة أنه احتاج إلى مراجعة. أما ١ فتعني أنه أخفق.
المعيار هو العمل المهني. تُقارن كل نتيجة بما كان سيقدمه محترف يتقاضى أجرًا لقاء المهمة نفسها.
تبقى المهام خاصة كي لا تتسرب إلى بيانات التدريب فتضخّم الدرجات مستقبلًا.
الهدف ليس إضافة قائمة ترتيب أخرى، بل مساعدتك في اختيار النموذج المناسب لنوع عملك. قد يتصدر نموذج في الاستدلال ويتأخر في الكتابة أو الذائقة التصميمية، وتُظهر الرسوم الرادارية موضع تفوقه وتأخره بدقة.
هذا الإصدار ٠٫١. وسيواصل التطور بمهام أصعب وأكثر صلة بالسوق، ومع صدور نماذج جديدة. لم تُدرج بعض النماذج الشائعة (أوبوس ٥، جي بي تي لونا) بعد، لأنني لم أجر عليها مهام كافية لتقييمها بإنصاف.
داحم بنش مستقل تمامًا. لا رعاة ولا علاقات مع الشركات المنتجة. دفعت تكلفة كل تشغيل من مالي الخاص، آلاف الدولارات حتى الآن. وهذا هو الهدف كله: نظرة صادقة ومحايدة إلى ما تفعله هذه النماذج فعلًا في العمل الحقيقي.
- Models profiledنموذجًا في التقييم
- 11
- Axes in four pillarsمحورًا في أربعة أركان
- 12
- Version, updated 20 September 2026الإصدار، حُدّث في 20 سبتمبر 2026
- v0.1
Rankingالترتيب
Overall score out of 10. Select a model to open its profile.المجموع من 10. اختر نموذجًا لفتح ملفه.
| Rankالمرتبة | Modelالنموذج | Overallالمجموع |
|---|
Model profilesملفات النماذج
Each axis is scored from 1 to 5. Add a second model to compare the two shapes.كل محور من 1 إلى 5. أضف نموذجًا ثانيًا لمقارنة الشكلين.
How the scoring worksكيف يُحسب التقييم
Every score is the evaluator’s own judgment from repeated use across several tools and contexts. It is a map of opinion, not a test suite.كل درجة حكم شخصي من استخدام متكرر في أدوات وسياقات متعددة. هذه خريطة رأي، لا مجموعة اختبارات.
- Scaleالمقياس
- 1 is weak, 5 is excellent, on every axis.من 1 ضعيف إلى 5 ممتاز، في كل محور.
- Overallالمجموع
- The mean of the scored axes, times two, out of 10.متوسط المحاور المقيّمة مضروبًا في اثنين، من 10.
- Not assessedغير مقيّم
- Left blank on the chart and never counted as zero.يُترك فارغًا في الرسم ولا يُحسب صفرًا.
- Updatesالتحديث
- Scores are revised as new experience builds up.تُراجع الدرجات كلما تراكمت تجربة جديدة.