najdarenaRequest evaluation
Najd Arena · By Najd Research

AI performance.
Measured in context.

Choosing a model starts with the problems it needs to solve. Najd Arena makes Arabic and Saudi AI performance easier to compare, understand, and reproduce.

Our purpose

Set a clearer standard
for relevant AI evaluation.

General benchmarks offer part of the picture. Arabic language, local knowledge, and business workflows deserve their own evidence.

We are building a reference for people choosing models for Arabic customer support, Saudi knowledge questions, document assistants, and tool-using agents.

Evaluation coverage

What we measure

Read the methodology →

Evidence you can inspect

Comparable settings

Direct model calls and agent-harness runs stay separate. Thinking settings remain visible.

Transparent scoring

Inspect task scores, counts, and failure handling. A small score difference is not a claim of statistical significance.

Reproducible foundations

Dataset versions and evaluation methods provide the context needed to interpret each result.

The currently imported results come from a preserved historical evaluation. They are not Najd-reviewed runs on the exact current dataset. Each result retains its scope and limitations.

Models in the current results

Provider logos identify evaluated models; they do not imply partnership or endorsement.

Founder-led research

Mazen Alotaibi

Founder, Najd Research.

Najd Arena is part of Najd Research. Follow the research, explore the datasets, or help define evaluations that reflect the work you need AI to do.

Explore Najd Research ↗
Build with us

Bring a model.
Bring a hard problem.

Discuss a benchmark, contribute to the work, or report a result that needs correction.

salam@najdresearch.com ↗