Comparable settings
Direct model calls and agent-harness runs stay separate. Thinking settings remain visible.
Choosing a model starts with the problems it needs to solve. Najd Arena makes Arabic and Saudi AI performance easier to compare, understand, and reproduce.
General benchmarks offer part of the picture. Arabic language, local knowledge, and business workflows deserve their own evidence.
We are building a reference for people choosing models for Arabic customer support, Saudi knowledge questions, document assistants, and tool-using agents.
Arabic understanding, Saudi knowledge, and culturally relevant questions.
Available results ↗02Answering questions from supplied information, with task-level breakdowns.
Available text results ↗03Models inside Pi on the available coding tasks, keeping execution separate from direct model calls.
Limited task coverage ↗04Visual understanding, Arabic OCR, document extraction, charts, and tables.
In development ↗Direct model calls and agent-harness runs stay separate. Thinking settings remain visible.
Inspect task scores, counts, and failure handling. A small score difference is not a claim of statistical significance.
Dataset versions and evaluation methods provide the context needed to interpret each result.
The currently imported results come from a preserved historical evaluation. They are not Najd-reviewed runs on the exact current dataset. Each result retains its scope and limitations.
Provider logos identify evaluated models; they do not imply partnership or endorsement.
Founder, Najd Research.
Najd Arena is part of Najd Research. Follow the research, explore the datasets, or help define evaluations that reflect the work you need AI to do.
Explore Najd Research ↗Discuss a benchmark, contribute to the work, or report a result that needs correction.