Nº IPositioning

Tell us your workload and required quality.

We show you the cheapest model likely to meet it.

Raw comparison is already crowded. This is not another universal leaderboard: pick a workload, set the minimum capability you will accept, and the cheapest models that clear it are ranked by what a successful outcome actually costs. Nothing here blends cost and intelligence into one opaque score.

The wedge
01Tokenizer-normalized pricing
02Workload-specific capability thresholds
03Actual cost per successful outcome
04Transparent, reproducible evidence
IIThe instrument10 of 14 models qualify

Select a workload and minimum capability, then rank the cheapest that meet it.

Capability is workload-specific. Coding, extraction, multilingual writing and difficult reasoning are evaluated separately, never collapsed into one property.

Prototype evidence boundary: model identities, rates, prices and latency below are illustrative interface fixtures, not measured evaluations or current provider evidence.

Workload
95.0%
80%99.5%
Models below the bar are excluded, not penalised
Input / output mix
Cache hit rate
Max latency p50

10 models qualify for Structured data extraction at 95.0%.

Ranked by total API cost ÷ successful outcomes · Structured data extraction · threshold 95.0%1,200 attempts per model · validator v4 · eval 12 Aug 2026
ModelValid-result rate$ / 1M SLUAttempts → passesCost per successp50
Excluded — 4 below the bar
GPT-4.1 mini · 94.8%Gemini 2.5 Flash-Lite · 92.8%Mistral Small 3.2 · 92.3%Amazon Nova Lite · 91.7%
Cost per successful outcome
total actual API cost across attemptsnumber of passing outcomes

A cheap attempt can still be expensive if it rarely produces a valid result.

IIIThe frontiercost versus capability · Structured data extraction

Above the line is eligible. On the curve is undominated.

Models inside the curve are beaten on both axes at once—cheaper alternatives exist at equal or better capability. Drag the threshold and the frontier moves.

88%91%94%97%100%Threshold 95.0%Claude Haiku 4.5GPT-5DeepSeek-V3.2Qwen3 235BCost per successful outcome — log scale · USD
IVModel dossierclick any row or chart point to load it here
Anthropic

Claude Haiku 4.5

List in / out$1.00 / $5.00
Context200k
p50 latency640 ms
Category results — cost per success by workload
Structured data extraction95.7%$0.00625
Coding94.6%$0.00632
Difficult reasoning93.4%$0.00640
Multilingual writing94.5%$0.00633
Conversation97.5%$0.00613
Structured output95.4%$0.00627
Price history — $ / 1M SLU, blendedNov 25Feb 26May 26Aug 26$2.20Evidence
Attempts per category1,200
Valid-result rate95.7% ±1.1
Evaluation date12 Aug 2026
Price effective date15 Oct 2025
Attempt shape3,200 in / 700 out SLU
VHandheld390 × 844

One question per screen.

On a phone the control bar becomes a sheet and the table becomes a stack. The threshold stays visible at all times, because it is the only input that changes the answer entirely.

9:41Threshold
Workload
Structured data extractionChange
Min rate95.0%
1Qwen3 235B95.2% pass · $0.29 per 1M$0.0010
2DeepSeek-V3.296.7% pass · $0.28 per 1M$0.0010
3Llama 4 Maverick95.2% pass · $0.40 per 1M$0.0013
4GPT-5 mini96.4% pass · $0.74 per 1M$0.0021
4 models excluded below the bar
Rankings

Pick a workload

Nº VIIndependence

Which model is cheapest for my workload while still meeting my required quality?

Evaluations are repeated, graded by published rules, and priced from provider-reported consumption. Every ranking carries its confidence interval, evaluation date and the effective date of its prices. Providers may pay for expedited testing or certified reports. Payment can never affect a score or rank.

What this site will not do
Advertising-driven rankings
Providers paying for higher placement
One unexplained ‘best model’ score
Affiliate revenue that can influence ordering
A gateway before the dataset has earned trust
Nº VIIPlans · provisional

The free product is the acquisition channel, not the business.

Stage 1

Free public product

£0

· Rankings and model pages

· Basic calculator

· Full methodology

Stage 2

Developer Pro

£19–39 / month · test price

· Saved workload profiles

· Custom weights and mixes

· Price and ranking alerts

· CSV and JSON export

· API allowance

Stage 3

Team evaluation

Quoted · from one-off studies

· Private model comparisons

· Cost-per-success analysis

· Procurement-ready reports

· Scheduled re-evaluation

· SSO and audit logs

Stage 4

Commercial data API

Licensed dataset

· Prices and effective dates

· Normalized SLU costs

· Capability results

· Historical ranking changes