Loading evaluation results…
Loading evaluation results…
Capability scores, comparable configurations, and the evidence behind each result.
Direct model calls · default thinking · Acceptable answer rate
Retrieval (RAG)
Alignment
Task highlights exclude subsets with fewer than 30 cases. Scores reflect different tasks and rubrics; they identify areas to inspect, not universal strengths or weaknesses. Input / output: text / text. Provider specifications are listed below when available.
Historical results Acceptable output rate · Higher is better
Historical results · Evaluation and publication dates have not been verified. Evidence policy →
Fixed 0–100% scale · Failures count as zero.
Raw is direct model execution; Pi uses the agent runtime. These paths and their prompt settings are kept separate. Levels are API settings, not verified measures of thinking effort. “Default” uses provider defaults; “disabled” requests no reasoning. Settings may behave differently across providers.
Each overall configuration covers 6,089 case IDs. These historical pre-audit results include all cases. Earlier case contents differ from the current release; no significance test is available. Hover, focus, or tap a bar to inspect its counts.
Choose a problem area. Task scores use Direct · default thinking.
See the full performance profile before narrowing your choice. Strong results in one area can hide weaknesses in another; use individual tasks to match the model to your workload.
Problem areas are navigation groups, not new composite indexes. Bars use a fixed 0–100% scale. Repeated configurations do not add independent cases.
Every configuration covers 6,089 case IDs. These are direct model calls. Keep the thinking setting fixed when comparing models.
| Prompt | Reasoning | Acceptable rate | Technical failures |
|---|---|---|---|
| raw | default | 66.12% | 10 |
| raw | disabled | 61.83% | 0 |
| raw | high | 66.20% | 9 |
| raw | low | 66.18% | 10 |
| raw | medium | 65.97% | 15 |
| raw | minimal | 65.77% | 9 |
| raw | xhigh | 65.79% | 12 |
Controlled cost and speed comparisons require endpoint-specific measurements under a declared workload. Partial elapsed-time records are available in the Inference page; they are not a standardized inference benchmark.
Acceptable = correct + possible correct, divided by all outputs. Technical failures contribute zero. All 6,089 cases are included, including quarantined cases.
Historical pre-audit experiment on 6,089 case IDs, not an exact run of the published Hugging Face case contents. Source prompts and expected fields differ on 1,419 and 1,422 cases respectively. Technical failures count as not acceptable in the full-denominator headline. No raw answers are public.
1,419 prompts and 1,422 expected answers differ from the published dataset. No uncertainty interval has been computed; small score differences do not establish a statistically significant advantage.
Pinned dataset reference ↗Evaluation metadata is reported separately from provider specifications.
| Evaluated input / output | Text / text |
|---|---|
| Recorded reasoning settings | disabled, default, minimal, low, medium, high, xhigh |
| Recorded prompt settings | Direct model execution (raw). Pi harness results appear under Coding Agents. |
| Context window, parameters, release date | Not verified for this imported model identifier |
| API cost and speed | Not measured |