najdarenaRequest evaluation
Coding agents / Evaluation coverage

Coding Agents

Compare models running inside the Pi harness on coding tasks. The results below keep the harness fixed while you explore models and thinking levels.

The coding-question subset contains only five case IDs per configuration. It is not a coding-agent leaderboard or a broad measure of software engineering ability.

Coding performance with Pi

Historical results Acceptable output rate · Higher is better

Pi harness
default
HUMAIN M3Limited evidence · 5 cases5 / 5
MiniMax M3Limited evidence · 5 cases5 / 5

Historical results · Evaluation and publication dates have not been verified. Evidence policy →

Fixed 0–100% scale · Failures count as zero.

How to read this chart

Raw is direct model execution; Pi uses the agent runtime. These paths and their prompt settings are kept separate. Levels are API settings, not verified measures of thinking effort. “Default” uses provider defaults; “disabled” requests no reasoning. Settings may behave differently across providers.

Each overall configuration covers 6,089 case IDs. These historical pre-audit results include all cases. Earlier case contents differ from the current release; no significance test is available. Hover, focus, or tap a bar to inspect its counts.