Free public product
£0· Rankings and model pages
· Basic calculator
· Full methodology
Raw comparison is already crowded. This is not another universal leaderboard: pick a workload, set the minimum capability you will accept, and the cheapest models that clear it are ranked by what a successful outcome actually costs. Nothing here blends cost and intelligence into one opaque score.
Capability is workload-specific. Coding, extraction, multilingual writing and difficult reasoning are evaluated separately, never collapsed into one property.
Prototype evidence boundary: model identities, rates, prices and latency below are illustrative interface fixtures, not measured evaluations or current provider evidence.
10 models qualify for Structured data extraction at 95.0%.
A cheap attempt can still be expensive if it rarely produces a valid result.
Models inside the curve are beaten on both axes at once—cheaper alternatives exist at equal or better capability. Drag the threshold and the frontier moves.
On a phone the control bar becomes a sheet and the table becomes a stack. The threshold stays visible at all times, because it is the only input that changes the answer entirely.
Evaluations are repeated, graded by published rules, and priced from provider-reported consumption. Every ranking carries its confidence interval, evaluation date and the effective date of its prices. Providers may pay for expedited testing or certified reports. Payment can never affect a score or rank.
· Rankings and model pages
· Basic calculator
· Full methodology
· Saved workload profiles
· Custom weights and mixes
· Price and ranking alerts
· CSV and JSON export
· API allowance
· Private model comparisons
· Cost-per-success analysis
· Procurement-ready reports
· Scheduled re-evaluation
· SSO and audit logs
· Prices and effective dates
· Normalized SLU costs
· Capability results
· Historical ranking changes