najdarenaRequest evaluation
Model rankings

Leaderboards

Compare available model scores. Historical results and Najd-reviewed evaluations use separate protocols and are presented separately.

Rankings by benchmark

Historical results Acceptable output rate · Higher is better

default

Historical results · Evaluation and publication dates have not been verified. Evidence policy →

Fixed 0–100% scale · Failures count as zero.

How to read this chart

Raw is direct model execution; Pi uses the agent runtime. These paths and their prompt settings are kept separate. Levels are API settings, not verified measures of thinking effort. “Default” uses provider defaults; “disabled” requests no reasoning. Settings may behave differently across providers.

Each overall configuration covers 6,089 case IDs. These historical pre-audit results include all cases. Earlier case contents differ from the current release; no significance test is available. Hover, focus, or tap a bar to inspect its counts.

Najd-reviewed evaluations

Najd-reviewed results require all 5,717 eligible cases and use a macro-average across 21 tracks. The historical results above use a different dataset snapshot and scoring protocol.

Why leaderboards matter

Comparable evaluations help you build a shortlist without relying on marketing claims. Inspect the task breakdown and evaluation scope before choosing a model for your own workload.

Awaiting canonical results

No Najd-reviewed runs published yet

Imported historical results are available above. This separate section is reserved for runs completed under the current evaluation protocol.

Explore available model scores →
Compare evaluation protocols →