Latest model benchmark snapshot
Generated benchmark summaries are loaded from the public benchmark API when available.
Fallback rows are compiled into the site for resilience and are not live benchmark data.
| Provider | Model | Capability | Cost efficiency | Measured at |
|---|---|---|---|---|
| OpenAI | gpt-4.1-mini | — | — | Static fallback |
| Anthropic | claude-3.5-haiku | — | — | Static fallback |
| gemini-2.0-flash | — | — | Static fallback | |
| Groq | llama-3.3-70b | — | — | Static fallback |
Methodology
Composite of available public quality suites such as MMLU, HumanEval, HellaSwag, MBPP, and ARC when those generated scores exist. Missing capability data is shown as not enough data.
Latest generated cost-efficiency suite score, normalised to 0–100. Higher indicates stronger cost efficiency in the benchmark snapshot.
A weekly scheduled job can publish generated scores to the benchmark database. If the API or generated data is unavailable, this page shows a labelled static fallback only.
Benchmark data is provided for informational purposes. Actual performance varies by use case. Run evals on your own dataset using the ModelSpend evaluation framework.