najdarenaRequest evaluation
Najd Arena / Inference

Inference

Compare model-serving performance using measured endpoint results.

Why it matters

A useful answer must also arrive within your time and cost budget. Endpoint measurements help teams choose how to serve a model for interactive assistants, document processing, or high-volume workloads.

Recorded execution time

Exploratory timing from preserved runs · Direct (Raw) · default thinking

Partial coverage

Successful, graded records with a positive elapsed time only. Recovery passes and missing timings make this an observed sample, not a standardized speed ranking. It includes execution overhead; it is not time to first token or tokens per second.

HUMAIN M3
Median
4.32 sec
95th percentile
63.03 sec
3,794 timed outputs / 6,089 in this configuration
MiniMax M3
Median
11.55 sec
95th percentile
23.56 sec
3,736 timed outputs / 6,089 in this configuration
What the records do not measure

Provider inference cost, candidate token throughput, and time to first token cannot be reconstructed reliably. The usage fields in the grade file describe the judge, not the candidate model. No cost or token-speed estimate is displayed.

Output speed

Generation speed affects how long users wait for a complete answer.

Planned evaluation

Time to first token

The initial wait affects whether a conversation feels responsive.

Planned evaluation

Cost per task

Task cost helps teams budget for actual usage rather than token prices alone.

Planned evaluation

Controlled inference evaluations are coming

Rankings will appear when validated evaluations are available. Contribute an endpoint or discuss a benchmark with Najd.

Explore text model results → · Get in touch