Leaderboards
Compare available model scores. Historical results and Najd-reviewed evaluations use separate protocols and are presented separately.
Rankings by benchmark
Historical results Acceptable output rate · Higher is better
Historical results · Evaluation and publication dates have not been verified. Evidence policy →
Fixed 0–100% scale · Failures count as zero.
How to read this chart
Raw is direct model execution; Pi uses the agent runtime. These paths and their prompt settings are kept separate. Levels are API settings, not verified measures of thinking effort. “Default” uses provider defaults; “disabled” requests no reasoning. Settings may behave differently across providers.
Each overall configuration covers 6,089 case IDs. These historical pre-audit results include all cases. Earlier case contents differ from the current release; no significance test is available. Hover, focus, or tap a bar to inspect its counts.
Najd-reviewed evaluations
Why leaderboards matter
Comparable evaluations help you build a shortlist without relying on marketing claims. Inspect the task breakdown and evaluation scope before choosing a model for your own workload.
No Najd-reviewed runs published yet
Imported historical results are available above. This separate section is reserved for runs completed under the current evaluation protocol.
Explore available model scores →