najdarenaRequest evaluation
Research protocol

Evidence behind
every score.

Results are useful when you can see what was tested, how it was scored, and which comparisons are valid.

01 / AVAILABLE RESULTS

Imported model evaluations

The current model profiles use historical pre-audit results: 6,089 case IDs per configuration and 14 configurations per model.

Acceptable rate is correct plus possible correct, divided by all outputs. Quarantined cases remain included. Technical failures contribute zero.

These results used earlier case contents and are not Najd-reviewed evaluations of the current published dataset. Each model profile records the differences.

Explore measured results →
02 / CANONICAL PIPELINE

New endpoint evaluations

The current submission pipeline is configured for 5,717 eligible cases across 21 tracks. Its Najd score averages the track scores, rather than pooling every output.

Structured tasks use task-specific grading. Open-ended tasks use a pinned judge profile. Publication requires a complete run, grading coverage, explicit approval from an organization administrator, and approval from at least one Najd Arena administrator. The publication date records when the result becomes public.

“Najd-reviewed” means these checks and approvals were completed. It is not external certification or a guarantee of production suitability. This is a separate protocol from the imported results. Their scores must not be ranked together.

View canonical leaderboard →

How to read a comparison

Match the conditions

Compare the same dataset contents, task scope, prompt, and reasoning setting.

Inspect the denominator

Outputs across repeated configurations are not independent test cases.

Keep uncertainty visible

Current profiles have no confidence intervals. Small differences alone do not establish superiority.

Check what is missing

Cost, speed, and untested capabilities remain unmeasured until supporting evidence exists.

Sources & reproducibility

The public dataset is available on Hugging Face. Model profiles link its pinned revision and disclose historical content differences. Private evaluation material is not exposed in the interface.

Open dataset repository ↗