Coding Agents
Compare models running inside the Pi harness on coding tasks. The results below keep the harness fixed while you explore models and thinking levels.
Coding performance with Pi
Historical results Acceptable output rate · Higher is better
Historical results · Evaluation and publication dates have not been verified. Evidence policy →
Fixed 0–100% scale · Failures count as zero.
How to read this chart
Raw is direct model execution; Pi uses the agent runtime. These paths and their prompt settings are kept separate. Levels are API settings, not verified measures of thinking effort. “Default” uses provider defaults; “disabled” requests no reasoning. Settings may behave differently across providers.
Each overall configuration covers 6,089 case IDs. These historical pre-audit results include all cases. Earlier case contents differ from the current release; no significance test is available. Hover, focus, or tap a bar to inspect its counts.