najdarenaRequest evaluation
Evaluation by Najd Research

Know where your model works.
Know what to improve.

Evaluate your model or AI application on Arabic, Saudi knowledge, and the tasks your users depend on. Start with a focused question; agree on the evidence needed to answer it.

Private by defaultScope agreed before a runPublication requires your approval
Beyond a headline score

Build an evaluation around your decision.

Whether you’re comparing versions or investigating a weak area, we’ll discuss the tasks, measurements, and reporting your team needs.

01

Performance by task

Identify where results differ across Arabic language, local knowledge, retrieval, or tool decisions.

02

Failure analysis

Agree on an analysis of answer quality and failure patterns so your team can prioritize improvements.

03

A repeatable comparison

Define the dataset version, scoring, and execution settings for a meaningful comparison with future runs.

What can we discuss today?

Start with Arabic and Saudi text tasks, document/RAG assistants, or tool decisions. Custom scopes and agent workflows are agreed individually.

Vision, Arabic OCR, and speech are in development. You can register interest; availability is confirmed during scoping.

Read our measurement approach →
Start here

Tell us what you want to learn.

A few details help us propose a useful evaluation. No endpoint or API key is needed at this stage.

Keep this brief non-sensitive. Don’t include credentials, customer data, or private evaluation examples.

This form prepares an email; it does not submit or save your request.

Prefer to write directly? salam@najdresearch.com ↗

From question to evidence

A clear path, with you in control.

1

Agree on the scope

We discuss your use case, task coverage, success criteria, pricing, and timeline before work begins.

2

Evaluate privately

Arrange access securely, confirm execution settings, and review the results and limitations with your team.

3

Decide on publication

Results remain private unless your organization and at least one Najd Arena admin approve publication. Public results carry a publication date.

Before you request

Do I need an OpenAI-compatible endpoint?

An OpenAI-compatible endpoint is the intended integration path. For this first conversation, share only a model or application description. We’ll agree on access requirements before evaluation.

Will my results appear on the leaderboard?

Not automatically. Publication needs your organization’s approval and approval from at least one Najd Arena administrator. A private evaluation does not commit you to public results.

How much does an evaluation cost?

Pricing and timing depend on task coverage, model usage, and analysis depth. These are agreed during scoping. Sending a request does not start a paid run.

Can I submit an endpoint and run it myself?

Self-service endpoint submissions are not open yet. Start with the request above. If you already have access, visit your organization workspace.